sakuraboy9128-cmd/dsh-sensevoice-npu ↗★ 0

dsh-sensevoice-npu

在英特尔 NPU 上运行语音转文字服务 适合拥有英特尔 NPU 硬件并需要本地化、低延迟语音输入的用户。

套件
dsh-sensevoice-npu
相容性
待驗證
版本
0.2.0
授權
MIT
最近更新
2026年9月30日

安裝

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:sakuraboy9128-cmd/dsh-sensevoice-npu

Configuration

KeyDefaultMeaning
pythonPath%LOCALAPPDATA%\dsh-sensevoice-npu\venv\Scripts\python.exeworker interpreter (runtime/bootstrap.ps1 target)
dataRoot/speech-to-text/sensevoiceDSH speech root; models are read from /models
precisionint8int8 (model.int8.onnx) or fp32 (model.onnx)
deviceNPUOpenVINO device for the recogniser (NPU, GPU, CPU)
vadDeviceCPUSilero VAD device — tiny model, dynamic shapes, NPU rejects it
useVadtruetrim silence and split long speech into segments
buckets[64, 96, 128, 192, 256, 384, 512, 768, 1024] (plugin config sets [256])LFR-frame buckets the engine may compile
prewarmBuckets[256]buckets compiled before the worker reports ready
idleTimeoutMs600000release the worker after idling; compiled blobs stay cached

Each bucket costs roughly 800 MB on disk (baked IR + OpenVINO NPU cache) and one ~2.5 minute compile the first time. One bucket is enough for dictation: 256 frames ≈ 15 s of audio, and the NPU handles the padding cheaply.