--- name: ok-script-ocr-lang description: Maintain ok-end-field OCR matcher language JSON and OCR text-fix mappings. Use for assets/lang modules, active OCR locales, matcher references, or assets/ocr_fix/ocr_text_fix.json; GUI gettext belongs to ok-script-i18n. --- # OK Script OCR Language Resources `assets/lang/.json` supplies task OCR matchers and localized business data; `assets/ocr_fix/ocr_text_fix.json` extends matching for known OCR confusion. GUI text belongs to gettext (`$ok-script-i18n`). See `docs/dev/i18n_OCR配置流程.md` for the full project convention. ## Matcher schema and locales - Runtime reads one file per module, **not** `assets/lang//.json`. A top-level business key contains locale nodes; each node has exactly one of `{"string": "..."}`, `{"pattern": "..."}`, or `{"terms": [...]}`. The accessor returns a string, compiled regex, or list respectively. Do not mix types in one node. - `src/data/lang/__init__.py` defines `ACTIVE_LOCALES_CONFIG`. Currently `zh_CN` and `zh_TW` are active for task OCR. Unknown or inactive locales normalize to `zh_CN`; a missing locale node falls back to `zh_TW` for Traditional Chinese, otherwise `zh_CN`, then the first available node. - Missing files or keys yield an empty module or `None`. First check exact module/key spelling and locale before changing matcher code. ## Add or change a key 1. Edit `assets/lang/.json` with a semantic key (legacy `k_*` hash keys can remain). Supply both active OCR locales. Reference it as `self.lang..`. 2. Run `uv run --locked python -m unittest tests.TestCheckLang -v`. It checks code references in active locales and node shape, nonempty values, and regex syntax across all lang JSON nodes. Review the actual OCR region in a game window when behavior depends on visual matching. 3. Regenerate `src/data/lang/_lang_typed.py` with `uv run --locked python scripts/i18n/gen_lang_stubs.py` and inspect its diff. Tracked full-locale business text belongs in `assets/lang/`; maintain the six core locales `zh_CN`, `zh_TW`, `en_US`, `ja_JP`, `ko_KR`, and `es_ES`, preserving extra locale nodes too. Structured canonical data belongs in `assets/data/`. Keep local generated files at their existing paths and preserve their ignored status; do not change ignore rules to commit them. When adding a data-only module, update `DATA_ONLY_MODULES` in the stub generator; task OCR modules stay in the generated accessor hints. ## OCR confusion map `ocr_text_fix.json` maps complete misread text to correct text. The patch uses only **same-length** pairs to derive character substitutions and extends the caller's `match` after the framework's regex fix. It does not normalize OCR output or change `Box.name`. Use language patterns or business parsing for length-changing and word-order errors. The old `src/data/ocr_normalize_map.py` is gone. For a mismatch, confirm the runtime locale, key/node type, regex, search region, and generated hint before adding a confusion pair.