lovedheart/dsh-plugin-telegram ↗★ 2

dsh-plugin-telegram

DSH plugin for Telegram bot integration - send/receive messages via Telegram Bot API 适合需要通过Telegram Bot与DSH智能体进行交互和消息推送的用户。

パッケージ
dsh-plugin-telegram
互換性
未検証
Harness ピア範囲
^0.1.1-rc.1
Cordis ピア範囲
^4.0.1
バージョン
0.6.5
ライセンス
MIT
最終更新
2026/09/24

インストール

検証済み bundle がないか、互換性チェックに失敗しています。先にリポジトリの説明を読んでください。 README 全文を読む ↗

ドキュメント

README 全文を読む ↗

Configuration

OptionTypeDefaultDescription
botTokenstring""Bot Token。可留空,插件会按优先级查找:config → 环境变量 → DSH Credentials
baseUrlstring""Custom Telegram API base URL
allowedChatsstring[][]Allowed chat IDs (empty = all)
allowedUsersstring[][]Allowed user IDs (empty = all)
requireMentionbooleanfalseRequire @mention in groups
pollingEnabledbooleanfalseEnable long-polling
longPollTimeoutnumber30Polling timeout in seconds
defaultChatIdstring""Default chat ID for messages
maxMessageLengthnumber4000Max chars before splitting
parseModestring"HTML"Parse mode (HTML or Markdown)
injectToAgentbooleantrueInject messages to agent loop
agentResponseModestring"tool"Response mode: 'tool' or 'direct'
replyPrefixstring""Optional prefix for agent responses
directReplyTimeoutSecnumber3600(direct mode) Absolute safety cap (seconds) for the reply-forward watcher. The watcher is busy-aware — it follows the agent while it runs (long tool-call turns are fine) and forwards the reply the moment the agent goes idle with a fresh message; this cap only bounds pathological hangs. Short replies are still forwarded within seconds.
progressEnabledbooleantrueShow a live trajectory (tool calls + thinking) on Telegram while the agent works. Works in both direct and tool response modes.
progressDelaySecnumber5Only post the trajectory if the turn is still running after this many seconds (short turns show nothing).
progressIntervalMsnumber5000Minimum gap between in-place edits (Telegram rate-limits edits to ~1/s per message; lowered frequency cuts API load).
progressTrailLinesnumber3Streaming footer: how many recent activity lines (💭/🔧) to show under the in-place reply while the model is working (0 = off). Keeps the message visibly moving during long tool-call/reasoning stretches.
progressPerBlockCharsnumber240Max chars per trajectory line (a reasoning block or a tool call).
progressMaxCharsnumber1500Max chars of the whole trajectory message (tail-truncated, so the newest items survive).
progressTimeoutSecnumber3600Absolute cap before the trajectory self-cleans (pathological hangs only).
approvalEnabledbooleantrueWhen the agent's permission policy is ask and a tool call needs a decision (e.g. a sandbox escalation), post an inline-keyboard approval card (✅ 批准 / 🔁 一直允许 / ❌ 拒绝) to the owning chat instead of failing closed. See "Tool-guard approval" below.
approvalTimeoutSecnumber1800How long an approval card waits for a tap before expiring (cancelled). 0 = no expiry.
approvalForDefaultAgentbooleantrueAlso surface asks from the deployment's shared default agent to the phone. Before /new, a plain Telegram message routes to that agent, so this is what makes the card appear in the state you usually test in. Set false to limit cards to agents this plugin explicitly created (telegram-*). Requires defaultChatId.
approvalAlwaysPathstring''File where "🔁 一直允许" remembers are persisted (defaults to $DSH_HOME/telegram-approval-always.json). Set an absolute path to relocate.
questionsEnabledbooleantrueWhen the agent calls ask_user_question (pick an option / type your own), post an inline-keyboard question card to the owning chat and answer it right there (in-process waterfall answerer), so a phone-only user isn't left waiting on the browser. See "Question cards" below.
questionsForDefaultAgentbooleantrueAlso surface questions from the deployment's shared default agent to the phone (mirrors approvalForDefaultAgent). Set false to limit cards to agents this plugin explicitly created (telegram-*). Requires defaultChatId.
autopilotEnabledbooleantrueWhether the /autopilot command is available. Set false to disable full-auto mode entirely. See "Autopilot (full-auto mode)" below.
autopilotSandboxModestringdanger-full-accessSandbox mode appended to the session while a chat is in autopilot (the "global write" half). Defaults to full disk access.
autopilotWindowMsnumber10000How long an autopilot ask_user_question notice waits before auto-committing the recommended option (0 = commit immediately). Gives you a window to tap ✋ 接管 to take over.
questionsTimeoutSecnumber1800How long a question card waits for an answer before auto-cancelling (agent turn unblocks). 0 = no expiry. Mirrors approvalTimeoutSec.
sttEndpointstringhttp://127.0.0.1:18068OpenAI-compatible Whisper base URL used to transcribe inbound voice notes (same service dsh-tool-audio's transcribe_audio hits).
voiceTranscribebooleantrueWhen the user sends a voice note, transcribe it and reply with the text under the voice bubble (🎧). Requires forwardInboundMedia.
voiceTranscribeLanguagestringautoForce a language code (e.g. zh/en) for transcription, or auto to let Whisper detect it.
voiceTranscriptToAgentbooleantrueAlso include the transcript in the message injected to the agent, so it already has the words and does NOT re-run transcribe_audio.
subagentBoardEnabledbooleantrueWhile a session spawns subagents, keep ONE pinned message per chat showing each subagent live (task + status, and what it is doing). See "Live subagent board" below.
subagentBoardPinbooleantruePin the board message so it stays at the top of the chat (the "fixed" part).
subagentBoardRefreshMsnumber2000How often the board re-reads live child sessions and re-renders (edits are throttled to ~1.5 s regardless).
subagentBoardIncludeDescendantsbooleanfalseWhen true, also show nested subagents (a subagent that spawns another). Default shows only direct children.
subagentBoardMaxRowsnumber10Cap on subagents shown before the overflow collapses into a … 另有 K 个未显示 line (keeps the message under Telegram's 4096-char limit).
verbosebooleanfalseEnable debug and info logs (default: errors only)

Multi-Bot Configuration

Run multiple Telegram bots in one plugin instance. Add a top-level bots: array; every item is one bot. This is a v0.6.0 feature — the legacy single-bot config (top-level fields, no bots) keeps working unchanged (see "Backward compatibility" below).

Each bots[] item supports the same per-bot fields as the top-level config. An item field left unset falls back to the top-level value of the same field, so you only spell out what differs per bot.

FieldTypeDefault (when unset)Description
idstringautoBot id used for routing (see "id auto-generation"). Must be unique across bots (duplicates throw at startup).
tokenstringtop-level botTokenBot token. botToken is an accepted alias (either name works on an item).
envKeystring"TELEGRAM_BOT_TOKEN"Env var to read the token from (per-bot, so two bots read two different vars — no crosstalk).
credentialKeystringsame as envKeyDSH credentials key to fall back to for this bot's token.
baseUrlstringtop-level baseUrlPer-bot custom API base URL.
defaultChatIdstringtop-level defaultChatIdPer-bot default chat (used for card/approval routing and /new follow-ups).
allowedChatsstring[]top-level allowedChatsPer-bot allowed chat ids (empty = all).
allowedUsersstring[]top-level allowedUsersPer-bot allowed user ids (empty = all).
requireMentionbooleantop-level requireMentionPer-bot group @mention requirement (filtered against that bot's getMe).
injectToAgentbooleantop-level injectToAgentPer-bot message injection into the agent loop.
agentResponseModestringtop-level agentResponseModePer-bot 'tool' / 'direct' reply mode.

Other top-level fields (longPollTimeout, maxMessageLength, parseMode, pollingEnabled, replyPrefix, …) are also readable per item when set.

id auto-generation — when an item has no id:

  • it has a resolvable token → id = "bot-" + token.slice(0, 8);
  • otherwise (no token) → id = "default".
  • The legacy single-bot fallback (no bots field) always yields id "default".

Backward compatibility — omit bots (or set bots: [] / a YAML-coerced {}): the plugin builds one bot, id "default", entirely from the top-level fields. Existing single-bot cordis.yml files need no change and behave bit-identically.

Missing-token degradation — a bot whose token resolves to nothing (config + envKey + credentialKey all empty) is skipped with a warn log; it never throws. It is still registered (so clientFor/meFor report a precise "skipped" error) but gets no client/poller/command-menu. If all bots are skipped the plugin runs tools-only (tools still register; their !client guards return a helpful error when called) — the same "missing token → tools-only" semantics as the legacy path.

Per-bot token resolution priority (per item, in order): token (plain) → process.env[envKey] → DSH credentials envKey → DSH credentials credentialKey (the last one only when it differs from envKey). Using a distinct envKey per bot is the recommended way to avoid a second bot accidentally picking up the first bot's token.

Sending tools & telegram_get_info under multi-bot
  • Every send/edit/delete/media tool (telegram_send_message, telegram_send_photo, telegram_send_document, telegram_edit_message, telegram_delete_message, …) now accepts an optional bot parameter (a bot id). Resolution order: (1) explicit bot — must be a known, connected bot, else a clear tool error; (2) the owning bot of the target chat when exactly one bot routes that chatId (composite-key reverse lookup, the bot whose poller last routed a message for that chat); (3) when two bots share one chatId (e.g. both bots' private chats with the same owner), the calling agent decides (v0.6.5): sends go back out the bot that owns the agent running the tool call — not the first bot in the registry, which is what used to make bot B's photos arrive from bot A; (4) the first/legacy bot. A single-bot config is unaffected.
  • telegram_get_info now returns an ARRAY — one entry per configured bot: { id, username, botId, name, connected, defaultChat }. A legacy single-bot config returns exactly one entry (id: "default"), so single-bot callers see one item as before.
  • telegram_get_updates accepts a bot parameter; each bot keeps its own manual poll offset (independent per bot, and independent of that bot's background poller), so a manual poll on one bot never advances another's stream.
Two-bot example
- insert:
    - id: telegram
      name: '/path/to/dsh-plugin-telegram/lib/index.js'
      config:
        pollingEnabled: true
        longPollTimeout: 30
        # Top-level values act as the per-bot defaults (inherited when an item omits a field).
        requireMention: true
        agentResponseMode: 'direct'
        bots:
          - id: alice
            token: '123456...:AA...'        # or leave empty + envKey below
            defaultChatId: '100000000'
            allowedUsers: ['100000000']
          - id: bob
            envKey: 'TELEGRAM_BOT_TOKEN_BOB' # reads its own env var (no token on the item)
            defaultChatId: '200000000'
            allowedUsers: ['200000000']
Multi-bot caveats
  • chatId is NOT globally unique across bots — the same numeric chatId can exist under two bots. So all per-chat state (dedup, board, indicator, agent routing) is keyed by the composite key k(botId, chatId) ("botId::chatId"), never by bare chatId. Messages are therefore not cross-deduped across bots: the same (chatId, messageId) delivered through two different bots is processed once per bot.
  • Isolate tokens with envKey so a second bot does not inherit the first bot's TELEGRAM_BOT_TOKEN (the default envKey/credentialKey is the same TELEGRAM_BOT_TOKEN; point each extra bot at its own key).
  • getUpdates offset is per-bot (Telegram maintains one update stream per bot) — the plugin tracks a separate manual offset per bot for the telegram_get_updates tool, and each background poller tracks its own cursor (offset file telegram-poller-offset-.json).

Inbound voice transcription (🎧)

When the user sends a voice note, the plugin transcribes it via the local Whisper service (sttEndpoint, default the same proxy transcribe_audio uses) and, in the same step:

  1. Replies with the recognized text directly under the voice bubble — a reply to the voice message is the only way to show it "on the next line" (Telegram bots cannot edit another user's message). It is a quiet, plain-text message (🎧 …) so arbitrary recognized text can't trip the entity parser.
  2. Reuses that transcript in the note injected to the agent, so the agent already has the spoken words and does not need to call transcribe_audio again (saves a round-trip and lets the agent answer immediately).

Requires forwardInboundMedia: true (the file must be downloaded to transcribe). The transcript is shown verbatim — no LLM re-phrase — so display is fast and cost-free. The whole feature is best-effort: if the service is down or the audio is silent, the voice note is still injected and answered normally (just no transcript line). Set voiceTranscribe: false to turn it off.

Live trajectory (tool calls + thinking)

While the agent works on a Telegram message, a single editable message shows a rolling trail of its recent activity (in both direct and tool modes), plus a continuous "typing…" chat action. Modelled on QwenPaw's Telegram channel edit-in-place streaming. Each recent item is one line:

  • 💭 — a chunk of the model's thinking (streamed as it happens);
  • 🔧 : — a tool call (name + a compact argument preview).

The whole message is tail-truncated to progressMaxChars, so the newest items stay visible and the oldest scroll off (each line is separately capped at progressPerBlockChars). The final reply is not shown here — it is sent as its own message when the turn ends.

It is deleted the moment the turn ends (turn/end), self-cleans after progressTimeoutSec, and never shows for turns that finish before progressDelaySec. Set progressEnabled: false to turn it off. It is purely best-effort — a Telegram failure never affects the real reply.

Live subagent board (🧩)

When a session spawns subagents (via the subagent / subagent_fork tools), the plugin keeps a single pinned message per chat that shows, in real time, every subagent currently working — and each finished one, locked in place. Each subagent occupies at most two lines:

🧩 子代理看板 · 2 工作中 / 1 完成
🟢 重构认证模块 · 工作中
🔧 read src/auth/session.js
🟢 跑集成测试 · 工作中
正在启动…
✅ 整理依赖 · 已完成
已完成 · 用时 42s
  • Line 1 — status emoji + a short task name + status word (工作中 / 已完成 / …). The task name comes from the child's subagent/descriptor label (or the subagent tool call's description), truncated to fit.
  • Line 2 — what it is doing right now: the child's most recent tool call (🔧 name + args), else its latest reasoning, else 正在启动… until activity appears.
  • Real-time — a ticker r