chenkezhen480/dsh-semantic-memory ↗★ 1
dsh-plugin-semantic-memory
Semantic long-term memory for DeepSeek Harness: embedding-based retrieval over a persistent cross-session memory store, with model-facing tools, proactive per-question recall, automatic conversation summarization, and auto-selected embedding provider (apiKey present → API, otherwise local).
AI 分析
核心用途是提供基于向量检索的跨会话长期记忆。适合需要持久化记忆、自动摘要,且希望在本地(免 API 密钥)或云端灵活切换嵌入模型的用户。
インストール
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:chenkezhen480/dsh-semantic-memoryドキュメント
README 全文を読む ↗Usage
Provider selection
The embedding provider is chosen by mode (explicit deployment switch), falling
back to the automatic selection:
| Configuration | Provider |
|---|---|
mode: 'cloud' | API (OpenAI-compatible /embeddings endpoint); requires apiKey |
mode: 'local' | local (ONNX via @huggingface/transformers, offline), even with an apiKey set |
no mode, apiKey present (non-empty) | API |
no mode, no apiKey | local |
provider: 'local' (explicit) | local, even with an apiKey set |
provider: 'api' (explicit) | API; requires apiKey |
Switching deployment mode means editing mode in the in-package
cordis.patch.yml and restarting the Web profile with the same DSH launcher —
in-package config overrides outer layers (settings.yaml
or user profile patch rows only fill keys the package does not declare; they do
not override it). A restart is needed after patch-file changes; the settings
document (~/.dsh/settings.yaml, semantic-memory: section) hot-reloads for
the keys it is allowed to supply. The first local embed downloads the model
(~100 MB, cached in ~/.cache/huggingface; use remoteHost for a mirror).
Verify the plugin is live
Open a new session (existing sessions keep their original tool set) and ask
the model: "Do you have memory_ tools?"* — it should list memory_write,
memory_search, memory_forget, and memory_stats. The system prompt also
carries a ## Long-term memory section once memories exist.
What the model can do
- Persist on its own — state a durable preference, fact, or decision; the
injected guidance makes the model call
memory_writewithout being asked. - Ask it to remember — "记住:我在用硅基流动的 API" →
memory_write. - Recall — "我之前对回答风格有什么偏好?" → the per-turn semantic recall
surfaces relevant memories automatically;
memory_searchdigs deeper (supportskind,tags,workspace,limit,min_score). - Manage —
memory_forgetdeletes;memory_statssummarizes the store.
Automatic behaviors
| Trigger | Behavior |
|---|---|
| Every user message | Asynchronous embedding + search; the freshest per-session hits are injected into the next prompt assembly (## Long-term memory (recalled for your current question)) |
| Every N user messages (default 5) | The harness LLM distills only the messages since the last summary (per-session seq cursor — no re-digesting, nothing skipped) into memory entries, written with the auto tag; the cadence can be set with the DSH_SEMANTIC_MEMORY_SUMMARIZE_EVERY environment variable (0 disables, overrides the config document) |
| Prompt assembly, no fresh recall | Strongest memories (importance × recency × access) injected as fallback |
Where the data lives
- Store:
$DSH_HOME/memories/memories.jsonl(one JSON line per entry, vectors included; edit/backup freely). - Settings: in-package
cordis.patch.yml(deployment source of truth — package config overrides outer layers);~/.dsh/settings.yamlundersemantic-memory:only fills keys the package does not declare (hot-reloaded).
Troubleshooting
- No memory_ tools in a session* — the session predates the plugin; start a new one.
- First local embed is slow / fails — the model downloads on first use; set
remoteHost: https://hf-mirror.comin restricted networks. apiprovider errors — confirmmode/apiKeyare set andapiBasepoints at an OpenAI-compatible endpoint (a/v1base gets/embeddingsappended).- Auto-summary never fires — it needs the
llmandagentDefaultModelservices (present in the standard web profile) andautoSummarizeEvery > 0.
Configuration
| Key | Default | Meaning |
|---|---|---|
mode | (unset) | Deployment switch: local forces the local model, cloud forces the API (requires apiKey). Unset keeps the automatic selection. |
provider | auto | auto selects by apiKey (non-empty → api, else local); explicit local/api overrides. An explicit mode overrides both. |
localModel | Xenova/bge-small-zh-v1.5 | Local transformer model id. |
remoteHost | https://huggingface.co | Model download host; set https://hf-mirror.com in restricted networks. |
apiBase | https://api.siliconflow.cn/v1 | API base URL (an /embeddings route is appended). |
apiKey | '' | API key. When non-empty and provider is not explicitly local, the API provider is used. |
apiModel | BAAI/bge-m3 | API embedding model name. |
memoryPath | $DSH_HOME/memories/memories.jsonl | Store file path. |
promptTopK | 3 | Memories injected per system-prompt assembly (0 disables). |
maxSearchResults | 10 | Default memory_search hit cap. |
minScore | 0.35 | Default minimum relevance for search hits. |
halfLifeMs | 30 days | Memory strength half-life. |
autoSummarizeEvery | 5 | Auto-summarize every N user messages (0 disables; needs llm + agentDefaultModel). The DSH_SEMANTIC_MEMORY_SUMMARIZE_EVERY env var overrides this (0..100). |
summarizeWindow | 12 | Most recent messages included in one auto-summary. |
summarizeMaxTokens | 800 | Token budget for the summary call. |
summarizeTemperature | 0.2 | Sampling temperature for the summary call. |