moazzamak/dsh-voice-input ↗★ 0

dsh-voice-input

提供基于本地faster-whisper的离线语音输入功能 适合需要高隐私、免联网本地语音转文字输入草稿的用户。

套件
dsh-voice-input
相容性
待驗證
版本
0.1.0
授權
MIT
最近更新
2026年9月30日

同名套件的其他儲存庫

安裝

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:moazzamak/dsh-voice-input

Configuration

Every option has a working default, so configuration is optional. To change one, override the row in $DSH_HOME/profiles/web/cordis.patch.yml — a patch replaces a row's entire config, so restate every key you still want:

- id: voice-input
  name: dsh-voice-input
  config:
    model: small.en
    language: en
    computeType: int8
    timeoutMs: 300000
KeyDefaultMeaning
modelbase.enWhisper model size or name
languageenSpoken language, or auto to detect
computeTypeint8CTranslate2 compute type
timeoutMs300000Deadline for one transcription
polishconservativeClean the transcript with a model; off returns the raw text
polishTimeoutMs15000Deadline for the cleanup call alone
polishProvider / polishModelcurrent selectionPin cleanup to a specific route
pythonPathpackage .venvInterpreter with faster-whisper
scriptPathbundled CLIThe transcription script
cacheDir$DSH_HOME/cache/voice-modelsModel weights cache

The cleanup uses the model the deployment already selects, so it needs no second credential and no extra configuration. It is a normal ctx.llm.stream call — the same route the agent uses — and it runs with temperature: 0 under its own deadline.

Choosing a model

ModelSizeNotes
tiny.en~75 MBFastest, least accurate
base.en~150 MBDefault; roughly 1 s of CPU per 5 s of speech
small.en~500 MBNoticeably better, about 3× the compute

Drop the .en suffix for multilingual models (base, small) and pair them with language: auto. A model that is not cached downloads on first use and needs network access once.

If huggingface.co is blocked, set DSH_VOICE_HF_ENDPOINT to a mirror such as https://hf-mirror.com before starting DSH.