sakuraboy9128-cmd/dsh-sensevoice-npu ↗★ 0
dsh-sensevoice-npu
在英特尔 NPU 上运行语音转文字服务 适合拥有英特尔 NPU 硬件并需要本地化、低延迟语音输入的用户。
安裝
$
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:sakuraboy9128-cmd/dsh-sensevoice-npu說明文件
閱讀完整 README ↗Configuration
| Key | Default | Meaning |
|---|---|---|
pythonPath | %LOCALAPPDATA%\dsh-sensevoice-npu\venv\Scripts\python.exe | worker interpreter (runtime/bootstrap.ps1 target) |
dataRoot | /speech-to-text/sensevoice | DSH speech root; models are read from /models |
precision | int8 | int8 (model.int8.onnx) or fp32 (model.onnx) |
device | NPU | OpenVINO device for the recogniser (NPU, GPU, CPU) |
vadDevice | CPU | Silero VAD device — tiny model, dynamic shapes, NPU rejects it |
useVad | true | trim silence and split long speech into segments |
buckets | [64, 96, 128, 192, 256, 384, 512, 768, 1024] (plugin config sets [256]) | LFR-frame buckets the engine may compile |
prewarmBuckets | [256] | buckets compiled before the worker reports ready |
idleTimeoutMs | 600000 | release the worker after idling; compiled blobs stay cached |
Each bucket costs roughly 800 MB on disk (baked IR + OpenVINO NPU cache) and one ~2.5 minute compile the first time. One bucket is enough for dictation: 256 frames ≈ 15 s of audio, and the NPU handles the padding cheaply.