Select HarfBuzz's language system for a variant, script, region or extended language subtag, and for Chinese other than zh (#206) - #213
Merged
jakejackson1 merged 2 commits intoSep 17, 2026
Conversation
…tended language subtag, and for Chinese other than zh (#206) OtlTags::language() looked a tag up by its first subtag, with Chinese the one exception. HarfBuzz 14.3.1 reads more of it, and mPDF now tries the candidates in HarfBuzz's order (hb_ot_tags_from_language() and hb_ot_tags_from_complex_language()): 1. -fonnapa, -polyton, -arevmda, -provenc, -fonipa, -geok, -syre, -syrj or -syrn after any language 2. a retired tag read whole: art-lojban, i-hak, i-lux, i-navajo, no-bok, no-nyn, zh-min, zh-min-nan 3. ga-Latg IRT, mnw-TH MONT, ro-MD MOL then ROM 4. Chinese scripts and regions, for every Chinese language 5. an extended language subtag, zh-yue ZHH 6. the language subtag, trying each of its tags in turn Only subtags before the first singleton count, as in HarfBuzz. Ucdn::$ot_languages gains zh and the nineteen other Chinese languages HarfBuzz maps (yue ZHH, lzh ZHT, the rest ZHS), gives ga IRI then IRT, and adds nv as NAV then ATH. Hans makes yue and lzh ZHS; any other script or region leaves them their own tag, where zh and its other members take zh's rules. The five zh- region keys are removed: nothing has read them since #201, and zh-mo's ZHT was wrong. An extended language Ucdn has no key for keeps the language subtag's tags, so ar-afb stays ARA. HarfBuzz's table gives most extended languages their macrolanguage's tag and the rest their ISO 639-3 code; syncing the table is left to its own issue. A font offering only IRT now lays out lang="ga" with it, one offering MOL lays out ro-MD with it rather than ROM (DejaVu Sans and Serif offer both under latn), and NotoSans-Regular's NAV, IPPH and APPH now serve nv, -fonipa and -fonnapa. No golden master or snapshot sets a language, and no fixture moves. Noto-LanguageTags-Synthetic merges subsets of Noto Sans TC, Georgian, Myanmar and Syriac with a 'locl' lookup under each of PGR, IRT, MOL, MONT, PRO, SYRE, IPPH, KGE, ATH, ZHS, ZHT and ZHH. For every row of the issue's table mPDF now draws the glyph hb-shape --language draws, where before it drew the unsubstituted one or ZHS's. Test font: Noto-LanguageTags-Synthetic (Noto Sans TC 2.004, Noto Sans Georgian 2.005, Noto Sans Myanmar 2.107, Noto Sans Syriac 3.000, all OFL 1.1), built in fontTools 4.59.2. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e() whether a language is Chinese (#206) The singleton cut becomes one preg_replace, the language comes off the front of the subtags so one array serves every rule, chinese() owns the test for a Chinese tag, and the extended language is looked up once. An empty tag returns before any of it, as it did before #206. Every expectation still matches hb_ot_tags_from_script_and_language(), the 87,094-tag comparison with HarfBuzz's own table differs only where it did, and each mechanism's mutation fails the same tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
jakejackson1
force-pushed
the
fix/201-chinese-script-subtags
branch
from
September 17, 2026 01:47
035d522 to
b23641e
Compare
jakejackson1
force-pushed
the
fix/206-complex-language-tags
branch
from
September 17, 2026 01:47
105c142 to
86f0d6e
Compare
Member
Author
|
Rebased: #190 was squash-merged into |
jakejackson1
merged commit Sep 17, 2026
19da5b1
into
fix/201-chinese-script-subtags
27 checks passed
Member
Author
|
This PR merged into a base branch that had already been squash-merged, so its change never reached |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #206.
What changes
OtlTags::language()looked a tag up by its first subtag, and handled Chinese separately (#201). HarfBuzz 14.3.1 also reads variants, scripts, regions, extended language subtags, and languages with more than one tag. mPDF now builds its candidates in HarfBuzz's order, checks each against what the script offers, and falls back toDFLTas before.The order it implements, from
hb_ot_tags_from_language()(hb-ot-tag.cc) andhb_ot_tags_from_complex_language()(hb-ot-tag-table.hh):-fonnapaAPPH,-polytonPGR,-arevmdaHYE,-provencPRO,-fonipaIPPH,-geokKGE,-syre/-syrj/-syrnSYRE/SYRJ/SYRN. HarfBuzz'ssubtag_matches (p, limit, "-polyton", 8)searches every subtag after the language.strcmp):art-lojban,i-hak,i-lux,i-navajo(NAV, ATH),no-bok,no-nyn,zh-min,zh-min-nan. Without these, step 5 would readno-nynas Nkole.ga-LatgIRT (lang_matches, so the script must come straight after the language),mnw-THMONT,ro-MDMOL then ROM (subtag_matches, so the region can be anywhere).zh-yuegives ZHH. This runs after step 4, sozh-yue-HKgives ZHH by its region andzh-yue-Hansgives ZHH byyue, as in HarfBuzz.ot_languages2listsga→ IRI, IRT andnv→ NAV, ATH).As in
hb_ot_tags_from_script_and_language(), subtags after the first singleton (-x-,-u-, …) are ignored, and a tag starting withx-gives nothing.Chinese
chinese()now takes the language's own tag rather than assumingzh.Ucdn::$ot_languagesgainszh→ ZHS and the 19 other Chinese languages HarfBuzz lists:cdo cjy cmn cnp cpx csp czh czo gan hak hnm hsn luh mnp nan sjc wuu→ ZHS,yue→ ZHH,lzh→ ZHT. A language whose own tag is ZHS, ZHT or ZHH counts as Chinese. The rules follow HarfBuzz's generator (gen-tag-table.pyremovesyueandlzhfrom thezhmacrolanguage and gives each only a-Hansrule):-Hansgives ZHS for every Chinese language.zhand its ZHS members take all ofzh's rules:Hant-HKZHH,Hant-MOZHTM then ZHH,HantZHT, then regionhkZHH,moZHTM then ZHH,twZHT.yueandlzhkeep their own tag for any other script or region:yue-Hant-TWandyue-MOgive ZHH, andlzh-HKgives ZHT.hb-shapeconfirms this. The issue said every Chinese language takeszh's rules, but HarfBuzz doesn't do that for these two.Ucdn::$ot_languageszhand the Chinese languages above.gais now['IRI ', 'IRT '], andnv(which had no entry) is['NAV ', 'ATH ']. A value can now be an array.OtlTagsis the only reader in this fork, andOtl::_getOTLLangTag()is the only one upstream (checked withgh search code). Any third-party code reading the public static would see an array for these two keys.zh-cn,zh-hk,zh-mo,zh-sg,zh-twkeys. Nothing has read them since A Chinese script subtag never selects ZHS or ZHT, and bare zh and zh-MO pick different language systems from HarfBuzz #201,zh-mo's ZHT was wrong, andchinese()now owns every Chinese mapping.Where mPDF still differs from HarfBuzz
When HarfBuzz has no table entry for a three-letter code, it upper-cases the code and uses that as the tag. mPDF doesn't do this. So an extended language that
Ucdnhas no key for keeps the language subtag's tags:ar-afbstays ARA. That matches HarfBuzz for most extended languages, because its table gives them their macrolanguage's tag. Otherwise, adding step 5 would have turnedar-afbfrom ARA into DFLT. The same table gap means baremnwstill gives MON, where HarfBuzz gives MON then MONT. Both belong to #211.Scope: hand-coded rules, not a full table sync
Before choosing, I compared all 354 single-subtag keys in
Ucdn::$ot_languageswith HarfBuzz 14.3.1'sot_languages2/ot_languages3/ot_languages3_multi, including its ISO 639-3 fallback:mlMLR→MAL,hyHYE→HYE0,berBER→BBR,scsSLA→SCS)grcPGR→GRC,yidJII→YID,nsoSOT→NSO, …)UcdnA sync would change the first tag for 16 existing keys, and some of them (Malayalam, Armenian, Ancient Greek, Yiddish) affect real documents. It would also add about 800 keys. That's a different kind of change from this issue, so this PR hand-codes the issue's cases and their mechanisms, and #211 tracks the sync with these numbers and examples.
As a check that the mechanism matches HarfBuzz once the data does, I swapped HarfBuzz's own table into
Ucdn::$ot_languagesat runtime. I then compared the candidates for 87,094 generated tags (36 languages, crossed with scripts, regions, variants, extended languages and singletons) againsthb_ot_tags_from_script_and_language()called throughlibharfbuzz14.3.1. The only differences were HarfBuzz's ISO 639-3 fallback (deliberately not implemented) and 99 malformedi-<subtag>tags, which HarfBuzz rejects through a length check.Behaviour change
IRTbut notIRInow usesIRTforlang="ga". Before, it usedDFLT.lang="nv"now selects NAV, or ATH if NAV isn't offered.tests/data/ttforpackages/*/fontsoffers IRT, ATH, PGR, PRO, KGE, SYRE/J/N, MONT or HYE0. The ones that offer an affected tag:NotoSans-RegularandNotoSansMono-GDEF13-Subsetoffer NAV under latn, solang="nv"now uses it instead of DFLT.NotoSans-Regular,NotoSansMono-GDEF13-SubsetandCarlito-MarkAttachmentType-Subsetoffer IPPH, and the two Noto fonts also offer APPH.-fonipaand-fonnapanow select them.ro-MDnow uses MOL instead of ROM.Ucdnhas keys for are the Chinese ones. So the extended-language rule changes nothing outsidezh-*.Test font and tests
Noto-LanguageTags-Synthetic.ttf(4 KB) merges subsets of Noto Sans TC 2.004, Noto Sans Georgian 2.005, Noto Sans Myanmar 2.107 and Noto Sans Syriac 3.000 (all OFL 1.1) using fontTools 4.59.2, and replaces the GSUB. No script has features by default. Each language system has onelocllookup that substitutes a different glyph:a→hunder IPPH,iunder IRT,munder MOL,punder PRO,tunder ATHα→βunder PGRႠ→Ⴁunder KGEက→ခunder MONTܐ→ܒunder SYRE骨→一under ZHS,二under ZHT,三under ZHHNo script offers IRI, ROM or NAV, so
gaandnvshow the fallback to the second tag.LanguageTagLangSysTestchecks the glyph mPDF draws (viaTextRecordingMpdf) for every row of the issue's table and a few more. Each result was confirmed withhb-shape --language=… --unicodes=…:langhb-shapeel-polytonga-Latgro-MDmnw-THoc-provencsyr-Syreen-fonipaka-Geokyueyue-Hant-HKcmn-Hanslzhzh-yuezh-lzhganvcmn-Hant-TWyue-HansroOtlTagsTestadds:-x-or-u-, andx-aloneyue-Hant-TW,yue-MO,lzh-HK,lzh-Hans,cmn-MO,hak-HK,nanzh-yue-HK,zh-yue-Hans,zh-cmn-Hant, andar-afbkeeping ARAro-MDwith MOL → MOL; without MOL → ROM; with neither → DFLT. The same forgawith IRI/IRT, plusnv→ ATH andga-Latgwithout IRT → DFLTI ran every expectation (80 of them) through
hb_ot_tags_from_script_and_language()against the same offered tags, and all match.Mutation check, one mechanism at a time, running
OtlTagsTest,LanguageTagLangSysTestandChineseLangSysTest:el-GR-polytonzh-yue-Hanszhzh's rules applied toyue/lzhtooro-x-md,el-u-polytonFixtures
I did a cold regeneration: deleted
tests/Mpdf/tmp/mpdf,tmp/mpdfandtmp/ttfontdata, then rancomposer fontcache:update all,otldump:update all,shaping:update allandsubset:update all. The only files written were the four new ones forNoto-LanguageTags-Synthetic, and no existing fixture moved. The new otldump lists the six scripts with oneloclper language system. The shaping fixture shows no substitution, because the golden master sets no language.Verification
composer csclean.OtlTags.phporUcdn.php.Simplify pass
After the review, the second commit simplifies
languageSystems():preg_replacearray_shifted off, so a single subtag array serves every rulechinese()now decides for itself whether the language is ChineseTests, the HarfBuzz comparison and every mutation give the same results as before.
Suggestions I didn't take:
Ucdn::$ot_languagesis public and can be changed, so a cache could go stale, and the call costs microseconds.$retiredfrom language subtags: half of the entries would still need literals.Stack
This PR's base is
fix/201-chinese-script-subtags(#207), and the stack continues #207 → #202 → #200 → #198. #190, #188 and #186 have been squash-merged intogravitypdf. #207 is still open, so this branch is not rebased.Upstream mirror
Upstream, the equivalent is
Otl::_getOTLLangTag()(src/Otl.php:6150onmpdf/mpdfdevelopment) plusUcdn::$ot_languages. The mirror moveslanguageSystems(),chinese(),tags()and the$variants,$retired,$languageSubtagsand$chineseRegionstables intoOtlas private members, and makes the sameUcdnedits. It depends on the mirror of #191/#201, which introduced the ordered-candidates loop andchinese(). Apply it after that one.🤖 Generated with Claude Code