jryang1997/dsh-hold-to-dictate ↗★ 1
@jryang1997/dsh-hold-to-dictate
提供长按输入框语音听写转文字功能 适合需要通过长按手势或快捷键进行语音输入和听写转译的用户。
安装
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:jryang1997/dsh-hold-to-dictate说明文档
阅读完整 README ↗Usage
| Gesture | Result |
|---|---|
| Hover the input box | A hint appears in the tool row, between the mode chips and the model selector |
| Press and hold ≈0.3 s without moving | A ring draws at the pointer, then a capsule floats up above the card: a live dot, the level waveform, and one short word |
| Release | Transcript is inserted at the caret, not sent |
Press Esc while recording or transcribing | Cancel |
| Hold, then drag up ≥48 px or move the pointer off the box | The capsule and the card's outline wash red and the word changes to "release to discard"; release there to drop it, come back to keep it |
Hold Ctrl+Shift+Space, then release | The same gesture with no mouse: hold the chord to record, let go to transcribe |
| A plain click, a drag, or selecting text | Nothing happens; the arc retracts and the gesture never arms |
Keyboard
Ctrl+Shift+Space is the whole gesture without a pointer: hold it down to record, let go
to transcribe, Esc while still holding to discard. It works wherever you are in the app, so
you do not have to focus the input box first, and it runs the same state
machine as the mouse: the capsule, the waveform, the discard wash and the retained-transcript
chip all behave identically.
The chord is awkward to hit by accident, and it never claims a keystroke unless it starts a recording.
It is not advertised anywhere in the interface. This plugin renders into a floating layer, which has nowhere to put a focusable control. It does not need one: the official Voice input bundle already puts a microphone button next to Send. That button is your visible, focusable entry to dictation; this chord is the keyboard route to this gesture, which inserts at the caret without switching modes.
Settings
Settings → Plugins → Hold to talk opens the bundle's own page. It has five settings:
| Setting | Why it is yours to change |
|---|---|
| Hold duration (150–800 ms, default 300) | Hands differ more here than at any other threshold |
| Hold duration on a touch screen (250–1200 ms, default 450) | A finger rolls, and a long press is also how a touch screen selects a word |
| Motion: full / calm | Calm keeps the cross-fades and drops movement and scale — the same softening as prefers-reduced-motion, except you choose it |
| Hover hint on / off | Some want the reminder, some find it noise |
Keyboard shortcut: off, Ctrl+Shift+Space, Ctrl+Shift+D, Ctrl+Shift+M, Ctrl+Alt+Space | Shortcut conflicts are personal, which is why this one is a choice |
Not in there: colours, materials, radii, the drag-to-discard distances, and the waveform's physics. The speech provider and language are absent for a different reason: this plugin sends neither, so the Host's own Voice input settings govern, and a second copy here could only disagree with them.
There is no theme setting: the plugin has no theme of its own. It carries no colour and reads every value from the Host's tokens, so it follows light and dark automatically.
Settings are stored in this browser, not in your DSH profile. That keeps
the plugin free of any @deepseek-ai/dsh-* dependency (a wrong peer range makes DSH
skip the whole bundle, silently), and these are per-machine preferences anyway. The cost is
that they do not follow you to another machine.
When something goes wrong
A failure is not a toast. It stays on screen until you deal with it: it is the one message that is waiting on you, and it announces itself assertively.
| Failure | What the card offers |
|---|---|
| The Host could not transcribe — a network hiccup, a provider error, a model that fell over | Retry, which re-sends the recording you already made. You should not have to say the same sentence twice |
| Anything the same recording would fail again on — too long, no recorder, models not prepared, microphone refused | Dismiss only, plus a line telling you what to change |
Dismiss (or Esc) clears it. A recording that was too short is not a failure and
still dissolves on its own.
If the draft changed while recognition was running, the transcript is kept in a small lower-right chip; click it to insert at the current caret.
使用
| 手势 | 结果 |
|---|---|
| 鼠标移入输入框 | 工具行中间(模式按钮与模型选择器之间)出现一行提示 |
| 按住不动约 0.3 秒 | 指针处先画出一个进度环,随后输入框上方浮起一颗胶囊:一个实心状态点、实时波形、四个字 |
| 松开 | 转写文字插入光标处,不发送 |
录音中或识别中按 Esc | 取消 |
| 按住后上滑 ≥48 px,或把鼠标移出输入框 | 胶囊与输入框外框一起泛红、文字变成「松开丢弃」;在那里松手即丢弃,移回来则保留 |
| 单击、拖拽、拖选文字 | 什么都不发生;进度环自己退回,手势不会激活 |
按住 Ctrl+Shift+空格,然后松开 | 同一套手势,只是不用鼠标:按住和弦开始录音,松开即转写 |
键盘
Ctrl+Shift+空格 就是没有指针的整套手势:按住录音,松开转写,按住时按 Esc 丢弃。
它在你处于应用任何位置时都有效,不必先聚焦输入框,走的是同一套状态机:
胶囊、波形、丢弃红晕、保留转写的胶囊,行为完全一致。
和弦不容易误触,只在真正开始录音时才拦截按键。
它没有在界面上任何地方宣传。 本插件渲染进一个浮层,那里没有位置放可聚焦的 控件。它也不需要:官方语音输入 bundle 已经在发送键旁边放了一个麦克风按钮。那个按钮 才是你看得见、键盘到得了的语音入口;而这个和弦是通往本插件这套手势的键盘路径 —— 它不切换 模式,直接插到光标处。
设置
「设置 → 插件 → 按住说话」会打开本 bundle 自己的页面。里面一共五项:
| 设置项 | 为什么它该由你决定 |
|---|---|
| 按住时长(150–800 ms,默认 300) | 在所有阈值里,这一项最因人而异 |
| 触屏按住时长(250–1200 ms,默认 450) | 手指会滚,而且触摸屏上的长按同时也是选词的起手 |
| 动效:完整 / 精简 | 「精简」保留淡入淡出、去掉位移与缩放,和 prefers-reduced-motion 是同一种软化,只是一个由你选、一个由系统给 |
| 悬停提示:开 / 关 | 有人要这个提醒,有人觉得是噪声 |
键盘快捷键:关闭、Ctrl+Shift+空格、Ctrl+Shift+D、Ctrl+Shift+M、Ctrl+Alt+空格 | 快捷键冲突是私人的,这正是它值得做成选项的原因 |
不放进去的:颜色、材质、圆角、上滑丢弃的距离、波形的物理参数。语音服务与语言缺席是另一个 原因:本插件两个都不发送,由 Host 自己的语音输入设置说了算,这里再放一份只会和它打架。
这里没有主题设置,因为本插件没有自己的主题:它一个颜色都不自带,所有值都读宿主令牌, 明暗主题自动跟随。
设置存在这个浏览器里,不在你的 DSH profile 里。这让插件不依赖任何
@deepseek-ai/dsh-* 包(peer range 写错会让 DSH 静默跳过整个 bundle),而这些本来就是
每台机器的偏好。代价是换台机器不携带。
出错的时候
失败不是一条一闪而过的提示。它会停在屏幕上直到你处理,因为这类消息是在等你回应, 用的是主动播报。
| 失败 | 卡片提供什么 |
|---|---|
| Host 转写失败 —— 网络抖动、provider 报错、模型挂了 | 重试,把你已经录好的那段重新发一次。不该让你把同一句话说第二遍 |
| 同一段录音再试也还是会失败 —— 太长、无录音能力、模型没准备、麦克风被拒 | 只有「关闭」,外加一行告诉你去改什么 |
「关闭」或 Esc 清掉它。录得太短不算失败,仍会自己消散。
录音不会接管输入框。 胶囊浮在卡片上方,卡片本身不动,所以录音期间草稿一直 可读、可编辑。
如果识别期间草稿被改动过,转写结果会保留在右下角的小胶囊里,点一下即可插入到当前光标。