site: twenty new languages, a language menu, and guides on corrupt files and on CI - #159
Conversation
…or more languages The site had two languages and the header held one pill per other language, which does not survive twenty. The switcher is now a details element with a list of links, no script, and the links stay in the page for a crawler. A language carries a locale beside its tag. og:locale was written from the tag, so it said "pl" where Open Graph defines "pl_PL"; the tag is still what lang, hreflang and the content directory use, and the locale is what a shared link says. A language can say it reads right to left, which puts dir="rtl" on its pages, and the stylesheet moves to logical properties so the same rules serve both directions. Code stays left to right and isolated inside an RTL page. Tracking is switched off for Arabic, Hindi and Thai, where space between letters breaks the word. Each non-Latin script gets a stack of system faces first - no font is loaded, the promise on the page is unchanged. Every page also carries a content-language meta, which Bing reads in place of hreflang. The sitemap is as it was, and so is the decision of 2026-08-27 to keep the alternates in the head only. The render now refuses languages that would put a wrong value on a page: a tag that is not BCP 47, a locale without a region, a prefix or a page address that cannot be part of one, two languages or two pages sharing one, and more or fewer than one language at the root. Guards: the boundary that kept characters above ASCII on the Polish pages is read from the language files, so it follows every translation. A new guard holds every page to its own lang, dir and og:locale, to the full set of hreflang alternates with x-default, to alternates that point back, and to a title and description of its own. The first-page limit example is asked of every translation and a language with no entry for its words fails. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
…abels German is the first of nineteen languages added to the site. Its text was written as a catalogue of the English strings and built into the same files the other languages have, so the structure, the links and every code sample are the English ones and cannot drift. The header keeps the language menu at the end of the row and lets the links wrap inside their own space, instead of pushing the menu to a second line when the labels of a language are long. On a narrow screen it scrolls away rather than taking a third of the phone. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
Four more translations, written as catalogues of the English strings and built into the same files every language has. Each carries its own slugs in ASCII (/fr/prereglages/, /es/documentacion/, /pt-br/casos-de-uso/, /it/preset/), its own title and description for every page, and the English code samples and links unchanged. French gets the spacing its typography asks for before a colon, a question mark and an exclamation mark, added at build time and kept out of code. The boundary that keeps characters above ASCII off the English pages now lets the list of languages in the header through, and only that. A reader on an English page has to be shown the way to Español and Türkçe in the names those languages use, and a guard that forbade it would have made the menu say "Spanish". The list is cut out of a page before it is read, so a language name anywhere else in English text is still a leak. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
Six more translations in the Latin scripts, each with page addresses in its own language written in plain ASCII (/nl/documentatie/, /ro/presetari/, /cs/predvolby/, /tr/dokumantasyon/, /vi/tai-lieu/ and so on), its own title and description for every page, and the English code samples and links unchanged. Romanian is written with the comma-below letters and Turkish keeps its dotted and dotless i. Every title and description was held to the length the English and Polish pages already keep, so a search engine does not cut them. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
Six translations in non-Latin scripts. Their pages sit under the English addresses with the language as a prefix (/ru/docs/, /zh-hans/formats/, /ja/faq/), so no address is percent-encoded and every link in the sitemap reads the same in a terminal, a log and a message. Simplified Chinese is served under zh-hans with the tag zh-Hans and Open Graph locale zh_CN, Traditional Chinese under zh-hant with zh-Hant and zh_TW, and the two are written separately in the vocabulary each audience uses rather than converted from one another. Titles and descriptions were measured by width, counting a wide character as two, against the lengths the English and Polish pages keep. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
The last three languages of the list. Arabic is the first right to left page set: dir=rtl on the root element, the layout mirrored by logical properties, and every command and code block kept left to right. The two guards this change added to site_test.go moved into sitelanguages_test.go, which keeps site_test.go under the crowding band the test-shape guard watches. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
…n the changelog The French pages carried the non-breaking space before a colon and inside guillemets as the character itself, which the guard against invisible characters refuses in a tracked file. In the page fragments it is now the entity, and the plain-text strings of site.json use an ordinary space. checkLanguages nested three deep and tipped the branching ceiling by one, so the claim on a unique value moved into its own method. Japanese pages break lines between phrases and keep small kana off the start of a line. The changelog says the website is available in twenty-two languages. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
…AQ answer that was out of date Two pages that answer questions people search for and that nothing on the site answered. How to make a corrupt file for testing: why a file corrupted by hand is a poor test, what a damaged file is here, the table of damages, what the manifest says about one, a recipe that mixes healthy and broken files, and a test that reads the manifest. The table is made from the program. The registry now reaches the site as Facts.Damages, every language has a sentence for every damage, and a guard asks for both directions, so a second damage cannot be added without the page and its translations following. How to generate test files in a CI pipeline: why a fixture should not live in the repository, a recipe, a GitHub Actions workflow and a GitLab job whose install step was run against the v0.4.0 archive, the exit codes a pipeline meets, and the PowerShell catch. Both sit under the use cases page. A page may now name its parent in site.json, which keeps it out of the header and gives it the breadcrumb that the preset pages already had. The two are linked from the use cases and the documentation. The FAQ said a deliberately broken file was not possible yet. tfg damage has been in the tool since 0.3.0. The answer is rewritten in all twenty-two languages. Inline code inside a note may now break, because a whole command in a sentence is wider than a phone and pushed the page sideways. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
|
Important Review skippedToo many files! This PR contains 241 files, which is 141 over the limit of 100. To get a review, reduce the PR to 100 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry. ⚙️ Run configurationConfiguration used: Repository UI (base), Organization UI (inherited) Review profile: ASSERTIVE Plan: Advanced Run ID: ⛔ Files ignored due to path filters (333)
📒 Files selected for processing (241)
You can disable this status message by setting the
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary
The website goes from two languages to twenty-two, and gets two guides that answer questions nothing on the site answered. The window and the command line are not part of this.
New languages: Simplified and Traditional Chinese, Japanese, German, French, Spanish, Brazilian Portuguese, Italian, Indonesian, Russian, Turkish, Czech, Vietnamese, Hindi, Korean, Arabic, Romanian, Dutch, Ukrainian and Thai. Every page exists in every language, so the sitemap went from 26 to 330 addresses (15 pages in 22 languages).
What changes
Languages and search signals
lang,<meta http-equiv="content-language">,og:localeandog:locale:alternate, and the full set ofhreflangalternates withx-default, all in the<head>. The sitemap is unchanged, with noxhtml:link./de/dokumentation/). The others keep the English words under the language prefix (/ja/docs/).code, an Open Graphlocale, a nativename, a lower casedirand an optionalrtl. The render refuses a set of languages with a malformed or repeated value, instead of publishing a tag a search engine would quietly drop.<details>element with no script, listing each language by the name it calls itself.dir="rtl"is set on the root, and commands and code stay left to right. Fonts are named per script, and letter spacing is off for Arabic, Hindi and Thai.Two guides
parentinsite.json, which keeps it out of the header and gives it the breadcrumb the preset pages already had) and are linked from the use cases and the documentation.Facts.Damages, every language needs a sentence for every damage, and a guard checks both directions.A fix on the way
The FAQ said a deliberately broken file was not possible yet.
tfg damagehas been in the tool since 0.3.0. The answer is rewritten in all twenty-two languages.Guards
sitelanguages_test.go, which keepssite_test.gounder the size band the test shape guard watches.tfg damageprints. , because a tracked file may not carry an invisible character. The plain text strings insite.jsonuse an ordinary space.checkLanguageswas split so the branching ceiling stays where it was.How it was checked
CGO_ENABLED=0 go test -tags "$(cat .github/build-tags)" ./...passes except three tests:TestADirectoryWeCannotWriteInIsAPermissionProblem(the sandbox runs as root) and the twoTestEveryFormatSurvivesItsReferenceTooltests (LibreOffice rejects the Office files there). They fail the same way on an untouchedmainin the same environment.tfg verify). The two YAML samples parse. The GitLab job and the PowerShell snippet were not run on a real runner.Notes for the reviewer
TestThePresetsSayInTheWindowWhatTheySayOnTheSitecompares a preset's wording on the site with the window's, word for word, whenever the language tags match. A window translation into one of these languages has to use the same sentences as the site, or that guard turns red.lastmod. The generator has no honest date to put there.🤖 Generated with Claude Code
https://claude.ai/code/session_019j5uostnHRaiLNGqbJfa1t
Generated by Claude Code