KSF1216/chinese-script-policy ↗★ 0
chinese-script-policy
Harness-neutral Traditional Chinese enforcer + offline converter (skill, DSH bundle, CLI). One script axis either way + Cantonese/Japanese-only filters. Converts both ways; optional wording preference. Also: an LLM output guard, cp950/GBK console scan. 适合需要严格规范模型输出字体,或进行离线文本转换的用户。
Install
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:KSF1216/chinese-script-policyREADME
Read the full README ↗chinese-script-policy
Enforce the Chinese script you asked for — Traditional by default — on the Chinese you produce and store, and convert it fully offline.
GitHub · 中文說明 · Offline page · npm · CLI reference
One body of code, harness-neutral: the same repository is a DSH bundle, an Agent Skill (Claude Code and compatible harnesses), an npm package, a command-line tool, and a library for web applications. There is no separate "DSH build" and "generic build", so the two cannot drift apart.
The check is one script axis with two directions — use one at a time — plus two independent filter axes:
- Script axis (two directions, pick one): ask for Traditional and it reports Simplified-only glyphs (2,637); ask for Simplified and it reports Traditional-only glyphs (3,083). You name the target, and the tool reports what does not belong to it.
- Filter axes (each its own switch): Cantonese colloquial register, and Japanese-only characters and words (367 shinjitai and kokuji, plus 123 Japanese words). The Japanese axis catches text that looks Chinese but is Japanese — neither Traditional nor Simplified, so a Simplified-only table never sees it. Neither filter axis depends on which direction the script axis uses.
Conversion has three independently runnable steps: Traditional ↔ Simplified, Cantonese colloquial → written Chinese (only the parts that can never be written Chinese), and Japanese shinjitai → Chinese. All three run only when you ask for them.
An optional wording layer (--wording, one checkbox on the web page, off by default) decides whether a conversion also follows the target script's local vocabulary: Simplified → Traditional uses the Traditional preference table (軟件 → 軟體, 硬盤 → 硬碟); Traditional → Simplified uses the Simplified preference table (軟體 → 软件, 網路 → 网络). The direction picks the table, so the two cannot be mixed up. With the switch off you get pure glyph conversion (軟件 is correct Traditional on its own).
Zero runtime dependencies: the tables are compiled into checked-in JSON, so nothing pulls OpenCC in and dependencies stays empty. But "zero dependencies" does not mean "bare only" — the difference is when you use it and whether anything has to be installed first:
| When you want it | How | Install first? |
|---|---|---|
| Convert one document, or the machine has no Node and you would rather not open a terminal | Download dist/tradzh.html and double-click it (single file, offline) | No |
| You already have this directory (cloned, or installed as a skill) | node scripts/tradzh.js … directly | No |
| A permanent CLI, or use inside scripts and CI | npx chinese-script-policy …, or npm i -g chinese-script-policy (bins: tradzh, chinese-script) | Optional |
| DSH: write guard + GUI switches + skill registration | dsh plugin --profile web add chinese-script-policy | Yes |
| Claude Code / another harness, for the skill or the hook | Clone into that harness's skills directory, or take hooks.json | Clone only |
| A web application (server-side check or convert, or checking live in the browser) | Server: chinese-script-policy/lib; browser: dist/tradzh.html or bundle core | Server yes, browser no |
In other words, only the two "keep it running" paths need an install (the DSH plugin, and importing the library server-side). The single-file page and running the repo's own scripts need nothing, and npx fetches the package itself.
This is a tool for projects that must tell Traditional Chinese, Simplified Chinese, Cantonese colloquial writing and Japanese kanji apart. It does not claim that any one script is the correct one.
Highlights
- Report what does not belong to the script you asked for — on the command line, in the browser, or before a file is ever written
- Two independent filters on top: Cantonese colloquial register, and Japanese-only characters and words
- Convert Traditional ↔ Simplified, Cantonese → written Chinese, and Japanese shinjitai → Chinese, each step on its own
- One single-file offline page: no Node, no network, nothing uploaded, the original text never modified
- Zero runtime dependencies — the tables are checked-in JSON, so results do not change with the machine
- One implementation behind the CLI, the write hook, the DSH plugin and the web page, so they cannot disagree
- Reads UTF-8, BOM, UTF-16, Big5 and GB18030 by itself, and refuses to guess silently when it cannot tell
- Guards a local LLM's answers through a proxy, using that same decision function
Defaults
Checking — the script axis is one axis with two directions (you name the target; the tool reports what does not belong to it) and only one direction is ever on, next to two independent axes for Cantonese and Japanese:
| Front end | Script axis (pick one) | Cantonese / register axis | Japanese axis |
|---|---|---|---|
| Offline page | Two checkboxes: Traditional (report Simplified glyphs) ✅ on by default / Simplified (report Traditional glyphs) ⬜ off by default | ✅ on | ✅ on |
| CLI | --variant traditional (default, reports Simplified glyphs) / --variant simplified (reports Traditional glyphs) | needs --written | needs --japanese |
| Write hook (Claude Code format) | fixed to "ask for Traditional" (blocks Simplified glyphs) — it has no Simplified side | ✅ | ✅ |
| DSH plugin (GUI settings card) | one of three: ask for Traditional (default) / ask for Simplified / do not check | each can be turned off | each can be turned off |
⚠️ Never turn both directions on: feeding Traditional text (後面的軟件很乾淨) to the "ask for Simplified" side reports 3 glyphs (U+5F8C U+8EDF U+6DE8) — with both on, every Chinese document is blocked.
That is to say: the CLI checks only the Traditional direction of the script axis by default, and Cantonese and Japanese have to be asked for; the web page turns on the Traditional direction plus Cantonese plus Japanese (the Simplified direction stays off); the hook fixes the script axis to "ask for Traditional" and applies both filters with it; and the DSH plugin lets you pick one of the three on the settings card.
Conversion — the direction, and every step that can change meaning, are yours to ask for; only the two steps that leave the text itself unchanged are automatic:
| Step | Flag (CLI) | Other front ends | Default | Why |
|---|---|---|---|---|
| Script, Traditional ↔ Simplified | --to-traditional / --to-simplified | Web: two buttons; library: toTraditional / toSimplified | you pick a direction | The two directions produce completely different text; there is no sensible default |
| Register, Cantonese → written Chinese | --to-written | Web: one button; library: toWritten | off | Register is style; only the parts that can never appear in written Chinese are converted, and the rest is left to a model |
| Japanese, shinjitai → Chinese | --convert-japanese | Web: a "strip Japanese" checkbox; library: stripJapanese | off | A document may quote Japanese on purpose; a quotation should not be edited behind your back |
| Wording preference (local vocabulary) | --wording | Web: a "wording preference" checkbox; library: the 4th argument of toTraditional / toSimplified | off | 軟件 and 軟體 are both correct Chinese; swapping the word is a preference, not a correction. The direction picks the table (Simplified → Traditional uses the Traditional preference table, Traditional → Simplified the Simplified one), so they cannot be mixed up |
Glyph preference (裏面 → 裡面) | (no CLI flag) | automatic on all three (useTc is on by default) | always automatic | Two common ways to write the same character, collapsed to the common one; nothing about the meaning changes. To turn it off, call toTraditional(text, true, false) |
| Compatibility ideograph normalization | (no flag) | automatic on all three | always automatic | The text itself is unchanged, only the code point is standardized; there is nothing to ask about |
Install
Three shortest paths
# 1) Install nothing: download dist/tradzh.html and double-click it. Offline, no Node, no network.
# 2) Command line, nothing kept around (npx fetches it)
npx chinese-script-policy --dir .
npx chinese-script-policy --text "
" --to-traditional
# 3) DSH: one command installs the plugin (write guard + skill registration + GUI switches)
dsh plugin --profile web add chinese-script-policy
DSH
As a local skill (simplest) — put the directory under $DSH_HOME/skills/, and the directory name must be chinese-script-policy:
git clone https://github.com/KSF1216/chinese-script-policy.git "$env:USERPROFILE\.dsh\skills\chinese-script-policy"
A new session sees it (DSH watches the directory live; no restart needed).
As a DSH bundle — one plugin line, whose plugin registers SKILL.md into the skill registry with ctx.skills.register(...) (what DSH calls embedded skills), so nothing has to be copied into ~/.dsh/skills:
dsh plugin --profile web add chinese-script-policy # published on npm
dsh plugin --profile web add github:KSF1216/chinese-script-policy
dsh plugin --profile web add ./chinese-script-policy-1.2.0.tgz
Installing straight from GitHub makes pnpm ask for an
allowBuildsentry in that profile'spnpm-workspace.yaml(which amounts to allowing this package's code to run at install time). Use npm or the tarball if you would rather not be asked.
One install affects one profile.
webandheadlesseach have their owndsh.profile.bundles, so run it once per shell you want guarded (change--profile).headlesshas no GUI, so a settings card there is meaningless — use theconfigrow ofcordis.patch.ymlorsettings.yamlinstead; the guard itself works the same (the sametools/pre-execute). ⚠️ The settings card writes the global user layer (thechinese-script-policy:block in$DSH_HOME/settings.yaml, not per profile), so flipping a switch in the GUI changes every profile at once.
Other harnesses and the command line
SKILL.md is the generic Agent Skills format (YAML frontmatter plus Markdown), so any harness can read it directly:
git clone https://github.com/KSF1216/chinese-script-policy.git "$env:USERPROFILE\.claude\skills\chinese-script-policy"
git clone https://github.com/KSF1216/chinese-script-policy.git ./skills/chinese-script-policy
It also works as a command-line tool on its own (bins: tradzh, chinese-script):
node scripts\tradzh.js --dir . # check a whole directory tree
node scripts\tradzh.js --fix --to-traditional --write out.txt `. **No Node needed** |
| Front end **inside your own app** (check as you type) | `scripts/core.js` plus `scripts/*.json` | `import { createCore } from 'chinese-script-policy/core'` and inject the tables yourself (there is no fs) |
| Server on **Node** (Express / Next / Nuxt / Workers) | `scripts/lib.js` (reads its own tables) | `import { scanText, toTraditional, guardInspect } from 'chinese-script-policy/lib'` |
| Server **not on JS** (PHP / Python / Java) | `scripts/tradzh.js` | Call it as a subprocess, or hand the work to the browser page |
| **Do not use** | `index.mjs` (the DSH plugin), `lib/client.js` (the DSH settings card) | The latter touches `window` on import and throws `window is not defined` under Node |
```js
// server: check and convert (ESM and CJS both work)
import { guardInspect, toTraditional } from 'chinese-script-policy/lib';
const verdict = guardInspect(userText); // exactly the same decision as the DSH write hook
if (verdict) return res.status(422).json({ reason: verdict.reason });
const stored = toTraditional(userText); // conversion is a SEPARATE step; nothing is silently rewritten
// front end / Edge (no fs): the same engine, tables injected by you
import { createCore } from 'chinese-script-policy/core';
import simplifiedOnly from 'chinese-script-policy/tables/simplified-only' with { type: 'json' };
const core = createCore({ simplifiedOnly /* …the other eight */ });
⚠️ An ESM JSON import must carry
with { type: 'json' }, or Node throwsERR_IMPORT_ATTRIBUTE_MISSING(CJSrequire()does not need it). ⚠️core's named exports come from thescripts/core.mjsshim (the UMD wrapper hides them from Node's static analysis), andtest:apiasserts that the shim exports exactly the keys ofcore.js.
A working demo is included: examples/web-app/ is a zero-dependency Node HTTP server plus one front-end page, with guardInspect wired to the API boundary. Measured against the package as installed from npm:
| Request | Response |
|---|---|
GET / | 200 |
POST /api/check with Simplified text | 422 plus the raw reason (BLOCKED … U+8F6F U+51C0) |
POST /api/check with Traditional text | 200 {"clean":true} |
POST /api/convert to=traditional | 200 {"text":"後面的軟件很乾淨"} |
Write-time guard (three ways to install it)
The pre-write check is one decision (guardInspect / guardMessage in scripts/lib.js); the ways to install it differ only in how it is attached to the harness:
| Installation | Who it is for | Switches and configuration |
|---|---|---|
| DSH plugin (recommended) | DSH | The GUI settings card: enabled, one of three script choices, the register and Japanese switches, block versus warn, Windows script file types (.ps1 / .cmd, below) — applied the moment you save. The card is global (one change affects every profile); each profile needs its own install, and headless has no GUI, so its switches go in YAML |
| Claude Code / other harnesses | Any harness speaking the same hook protocol | Use this package's hooks.json (PreToolUse with matcher write|edit) |
| Environments without hooks | Everything else | Re-check yourself after writing with node scripts\tradzh.js , and put the policy in the system prompt |
Besides Chinese characters, Windows script file types carry two more write rules: .ps1 must be pure ASCII, and .cmd / .bat must be pure ASCII with CRLF — because PowerShell 5.1 and cmd.exe both read scripts as ANSI, and UTF-8 Chinese without a BOM turns a .ps1 into a syntax error. The rule is in SKILL.md under "file-type traps"; the cause and the measurements are in references/encoding.md. The **DSH plugin warns o