lovedheart/dsh-plugin-telegram ↗★ 2
dsh-plugin-telegram
DSH plugin for Telegram bot integration - send/receive messages via Telegram Bot API 适合需要通过Telegram Bot与DSH智能体进行交互和消息推送的用户。
Install
This plugin has no verified bundle, or compatibility checks failed. Read the repository notes first. Read the full README ↗
README
Read the full README ↗Configuration
| Option | Type | Default | Description |
|---|---|---|---|
botToken | string | "" | Bot Token。可留空,插件会按优先级查找:config → 环境变量 → DSH Credentials |
baseUrl | string | "" | Custom Telegram API base URL |
allowedChats | string[] | [] | Allowed chat IDs (empty = all) |
allowedUsers | string[] | [] | Allowed user IDs (empty = all) |
requireMention | boolean | false | Require @mention in groups |
pollingEnabled | boolean | false | Enable long-polling |
longPollTimeout | number | 30 | Polling timeout in seconds |
defaultChatId | string | "" | Default chat ID for messages |
maxMessageLength | number | 4000 | Max chars before splitting |
parseMode | string | "HTML" | Parse mode (HTML or Markdown) |
injectToAgent | boolean | true | Inject messages to agent loop |
agentResponseMode | string | "tool" | Response mode: 'tool' or 'direct' |
replyPrefix | string | "" | Optional prefix for agent responses |
directReplyTimeoutSec | number | 3600 | (direct mode) Absolute safety cap (seconds) for the reply-forward watcher. The watcher is busy-aware — it follows the agent while it runs (long tool-call turns are fine) and forwards the reply the moment the agent goes idle with a fresh message; this cap only bounds pathological hangs. Short replies are still forwarded within seconds. |
progressEnabled | boolean | true | Show a live trajectory (tool calls + thinking) on Telegram while the agent works. Works in both direct and tool response modes. |
progressDelaySec | number | 5 | Only post the trajectory if the turn is still running after this many seconds (short turns show nothing). |
progressIntervalMs | number | 5000 | Minimum gap between in-place edits (Telegram rate-limits edits to ~1/s per message; lowered frequency cuts API load). |
progressTrailLines | number | 3 | Streaming footer: how many recent activity lines (💭/🔧) to show under the in-place reply while the model is working (0 = off). Keeps the message visibly moving during long tool-call/reasoning stretches. |
progressPerBlockChars | number | 240 | Max chars per trajectory line (a reasoning block or a tool call). |
progressMaxChars | number | 1500 | Max chars of the whole trajectory message (tail-truncated, so the newest items survive). |
progressTimeoutSec | number | 3600 | Absolute cap before the trajectory self-cleans (pathological hangs only). |
approvalEnabled | boolean | true | When the agent's permission policy is ask and a tool call needs a decision (e.g. a sandbox escalation), post an inline-keyboard approval card (✅ 批准 / 🔁 一直允许 / ❌ 拒绝) to the owning chat instead of failing closed. See "Tool-guard approval" below. |
approvalTimeoutSec | number | 1800 | How long an approval card waits for a tap before expiring (cancelled). 0 = no expiry. |
approvalForDefaultAgent | boolean | true | Also surface asks from the deployment's shared default agent to the phone. Before /new, a plain Telegram message routes to that agent, so this is what makes the card appear in the state you usually test in. Set false to limit cards to agents this plugin explicitly created (telegram-*). Requires defaultChatId. |
approvalAlwaysPath | string | '' | File where "🔁 一直允许" remembers are persisted (defaults to $DSH_HOME/telegram-approval-always.json). Set an absolute path to relocate. |
questionsEnabled | boolean | true | When the agent calls ask_user_question (pick an option / type your own), post an inline-keyboard question card to the owning chat and answer it right there (in-process waterfall answerer), so a phone-only user isn't left waiting on the browser. See "Question cards" below. |
questionsForDefaultAgent | boolean | true | Also surface questions from the deployment's shared default agent to the phone (mirrors approvalForDefaultAgent). Set false to limit cards to agents this plugin explicitly created (telegram-*). Requires defaultChatId. |
autopilotEnabled | boolean | true | Whether the /autopilot command is available. Set false to disable full-auto mode entirely. See "Autopilot (full-auto mode)" below. |
autopilotSandboxMode | string | danger-full-access | Sandbox mode appended to the session while a chat is in autopilot (the "global write" half). Defaults to full disk access. |
autopilotWindowMs | number | 10000 | How long an autopilot ask_user_question notice waits before auto-committing the recommended option (0 = commit immediately). Gives you a window to tap ✋ 接管 to take over. |
questionsTimeoutSec | number | 1800 | How long a question card waits for an answer before auto-cancelling (agent turn unblocks). 0 = no expiry. Mirrors approvalTimeoutSec. |
sttEndpoint | string | http://127.0.0.1:18068 | OpenAI-compatible Whisper base URL used to transcribe inbound voice notes (same service dsh-tool-audio's transcribe_audio hits). |
voiceTranscribe | boolean | true | When the user sends a voice note, transcribe it and reply with the text under the voice bubble (🎧). Requires forwardInboundMedia. |
voiceTranscribeLanguage | string | auto | Force a language code (e.g. zh/en) for transcription, or auto to let Whisper detect it. |
voiceTranscriptToAgent | boolean | true | Also include the transcript in the message injected to the agent, so it already has the words and does NOT re-run transcribe_audio. |
subagentBoardEnabled | boolean | true | While a session spawns subagents, keep ONE pinned message per chat showing each subagent live (task + status, and what it is doing). See "Live subagent board" below. |
subagentBoardPin | boolean | true | Pin the board message so it stays at the top of the chat (the "fixed" part). |
subagentBoardRefreshMs | number | 2000 | How often the board re-reads live child sessions and re-renders (edits are throttled to ~1.5 s regardless). |
subagentBoardIncludeDescendants | boolean | false | When true, also show nested subagents (a subagent that spawns another). Default shows only direct children. |
subagentBoardMaxRows | number | 10 | Cap on subagents shown before the overflow collapses into a … 另有 K 个未显示 line (keeps the message under Telegram's 4096-char limit). |
verbose | boolean | false | Enable debug and info logs (default: errors only) |
Multi-Bot Configuration
Run multiple Telegram bots in one plugin instance. Add a top-level bots:
array; every item is one bot. This is a v0.6.0 feature — the legacy single-bot
config (top-level fields, no bots) keeps working unchanged (see "Backward
compatibility" below).
Each bots[] item supports the same per-bot fields as the top-level config. An
item field left unset falls back to the top-level value of the same field, so
you only spell out what differs per bot.
| Field | Type | Default (when unset) | Description |
|---|---|---|---|
id | string | auto | Bot id used for routing (see "id auto-generation"). Must be unique across bots (duplicates throw at startup). |
token | string | top-level botToken | Bot token. botToken is an accepted alias (either name works on an item). |
envKey | string | "TELEGRAM_BOT_TOKEN" | Env var to read the token from (per-bot, so two bots read two different vars — no crosstalk). |
credentialKey | string | same as envKey | DSH credentials key to fall back to for this bot's token. |
baseUrl | string | top-level baseUrl | Per-bot custom API base URL. |
defaultChatId | string | top-level defaultChatId | Per-bot default chat (used for card/approval routing and /new follow-ups). |
allowedChats | string[] | top-level allowedChats | Per-bot allowed chat ids (empty = all). |
allowedUsers | string[] | top-level allowedUsers | Per-bot allowed user ids (empty = all). |
requireMention | boolean | top-level requireMention | Per-bot group @mention requirement (filtered against that bot's getMe). |
injectToAgent | boolean | top-level injectToAgent | Per-bot message injection into the agent loop. |
agentResponseMode | string | top-level agentResponseMode | Per-bot 'tool' / 'direct' reply mode. |
Other top-level fields (longPollTimeout, maxMessageLength, parseMode,
pollingEnabled, replyPrefix, …) are also readable per item when set.
id auto-generation — when an item has no id:
- it has a resolvable
token→ id ="bot-" + token.slice(0, 8); - otherwise (no token) → id =
"default". - The legacy single-bot fallback (no
botsfield) always yields id"default".
Backward compatibility — omit bots (or set bots: [] / a YAML-coerced
{}): the plugin builds one bot, id "default", entirely from the
top-level fields. Existing single-bot cordis.yml files need no change and
behave bit-identically.
Missing-token degradation — a bot whose token resolves to nothing
(config + envKey + credentialKey all empty) is skipped with a warn log;
it never throws. It is still registered (so clientFor/meFor report a precise
"skipped" error) but gets no client/poller/command-menu. If all bots are
skipped the plugin runs tools-only (tools still register; their !client
guards return a helpful error when called) — the same "missing token → tools-only"
semantics as the legacy path.
Per-bot token resolution priority (per item, in order):
token (plain) → process.env[envKey] → DSH credentials envKey → DSH
credentials credentialKey (the last one only when it differs from envKey).
Using a distinct envKey per bot is the recommended way to avoid a second bot
accidentally picking up the first bot's token.
Sending tools & telegram_get_info under multi-bot
- Every send/edit/delete/media tool (
telegram_send_message,telegram_send_photo,telegram_send_document,telegram_edit_message,telegram_delete_message, …) now accepts an optionalbotparameter (a bot id). Resolution order: (1) explicitbot— must be a known, connected bot, else a clear tool error; (2) the owning bot of the target chat when exactly one bot routes that chatId (composite-key reverse lookup, the bot whose poller last routed a message for that chat); (3) when two bots share one chatId (e.g. both bots' private chats with the same owner), the calling agent decides (v0.6.5): sends go back out the bot that owns the agent running the tool call — not the first bot in the registry, which is what used to make bot B's photos arrive from bot A; (4) the first/legacy bot. A single-bot config is unaffected. telegram_get_infonow returns an ARRAY — one entry per configured bot:{ id, username, botId, name, connected, defaultChat }. A legacy single-bot config returns exactly one entry (id: "default"), so single-bot callers see one item as before.telegram_get_updatesaccepts abotparameter; each bot keeps its own manual poll offset (independent per bot, and independent of that bot's background poller), so a manual poll on one bot never advances another's stream.
Two-bot example
- insert:
- id: telegram
name: '/path/to/dsh-plugin-telegram/lib/index.js'
config:
pollingEnabled: true
longPollTimeout: 30
# Top-level values act as the per-bot defaults (inherited when an item omits a field).
requireMention: true
agentResponseMode: 'direct'
bots:
- id: alice
token: '123456...:AA...' # or leave empty + envKey below
defaultChatId: '100000000'
allowedUsers: ['100000000']
- id: bob
envKey: 'TELEGRAM_BOT_TOKEN_BOB' # reads its own env var (no token on the item)
defaultChatId: '200000000'
allowedUsers: ['200000000']
Multi-bot caveats
chatIdis NOT globally unique across bots — the same numericchatIdcan exist under two bots. So all per-chat state (dedup, board, indicator, agent routing) is keyed by the composite keyk(botId, chatId)("botId::chatId"), never by barechatId. Messages are therefore not cross-deduped across bots: the same(chatId, messageId)delivered through two different bots is processed once per bot.- Isolate tokens with
envKeyso a second bot does not inherit the first bot'sTELEGRAM_BOT_TOKEN(the defaultenvKey/credentialKeyis the sameTELEGRAM_BOT_TOKEN; point each extra bot at its own key). getUpdatesoffset is per-bot (Telegram maintains one update stream per bot) — the plugin tracks a separate manual offset per bot for thetelegram_get_updatestool, and each background poller tracks its own cursor (offset filetelegram-poller-offset-.json).
Inbound voice transcription (🎧)
When the user sends a voice note, the plugin transcribes it via the local
Whisper service (sttEndpoint, default the same proxy transcribe_audio uses)
and, in the same step:
- Replies with the recognized text directly under the voice bubble — a reply
to the voice message is the only way to show it "on the next line" (Telegram
bots cannot edit another user's message). It is a quiet, plain-text message
(
🎧 …) so arbitrary recognized text can't trip the entity parser. - Reuses that transcript in the note injected to the agent, so the agent
already has the spoken words and does not need to call
transcribe_audioagain (saves a round-trip and lets the agent answer immediately).
Requires forwardInboundMedia: true (the file must be downloaded to transcribe).
The transcript is shown verbatim — no LLM re-phrase — so display is fast and
cost-free. The whole feature is best-effort: if the service is down or the audio
is silent, the voice note is still injected and answered normally (just no
transcript line). Set voiceTranscribe: false to turn it off.
Live trajectory (tool calls + thinking)
While the agent works on a Telegram message, a single editable message shows a
rolling trail of its recent activity (in both direct and tool modes), plus a
continuous "typing…" chat action. Modelled on QwenPaw's Telegram channel edit-in-place
streaming. Each recent item is one line:
💭— a chunk of the model's thinking (streamed as it happens);🔧 :— a tool call (name + a compact argument preview).
The whole message is tail-truncated to progressMaxChars, so the newest items
stay visible and the oldest scroll off (each line is separately capped at
progressPerBlockChars). The final reply is not shown here — it is sent as its own
message when the turn ends.
It is deleted the moment the turn ends (turn/end), self-cleans after
progressTimeoutSec, and never shows for turns that finish before progressDelaySec.
Set progressEnabled: false to turn it off. It is purely best-effort — a Telegram
failure never affects the real reply.
Live subagent board (🧩)
When a session spawns subagents (via the subagent / subagent_fork tools), the
plugin keeps a single pinned message per chat that shows, in real time, every
subagent currently working — and each finished one, locked in place. Each subagent
occupies at most two lines:
🧩 子代理看板 · 2 工作中 / 1 完成
🟢 重构认证模块 · 工作中
🔧 read src/auth/session.js
🟢 跑集成测试 · 工作中
正在启动…
✅ 整理依赖 · 已完成
已完成 · 用时 42s
- Line 1 — status emoji + a short task name + status word (
工作中/已完成/ …). The task name comes from the child'ssubagent/descriptorlabel (or thesubagenttool call'sdescription), truncated to fit. - Line 2 — what it is doing right now: the child's most recent tool call
(
🔧 name + args), else its latest reasoning, else正在启动…until activity appears. - Real-time — a ticker r