ranxianglei/billion-context ↗★ 211

billion-context

Context-compression proxy that lets AI coding agents run for days — billions of tokens through one context window. Sits between any agent (Claude Code, Codex, Cursor, Aider, …) and its model API, folding consumed conversation into reversible, prefix-cache-friendly summaries. Zero per-agent code: just point your base URL at the proxy. 适合需要处理超长上下文、减少 Token 消耗的 AI 编码任务。

Package
billion-context
Compatibility
Unverified
Version
0.1.129
License
MIT
Last updated
Sep 20, 2026

Install

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:ranxianglei/billion-context

Configuration

The full configuration reference — config file location, top-level keys, providers, compression tuning, environment variables — lives in CONFIGURATION.md.

Upstream proxy (firewall / GFW)

If the proxy's own outbound connections to a model provider are blocked (e.g. api.openai.com from inside the GFW), configure an upstream proxy (the local v2rayA / clash HTTP port) so the proxy reaches the provider:

{
  // Global default: ALL providers route through this proxy
  "proxy": "http://127.0.0.1:20172",
  "providers": {
    "https://api.openai.com/v1": {
      // Per-URL overrides global (use a different proxy for this host)
      "proxy": "http://127.0.0.1:20173",
      "models": { "gpt-5": { "context": 400000 } }
    },
    "https://open.bigmodel.cn/api/anthropic": {
      // Empty string = explicitly DIRECT, overriding the global proxy
      "proxy": "",
      "models": { "glm-5.2": { "context": 1000000 } }
    }
  }
}

Rules:

  • Per-URL proxy has the highest priority for its matching provider URL.
  • Remaining priority is BILI_UPSTREAM_PROXY → Web UI manual proxy → top-level proxy → HTTPS_PROXY / HTTP_PROXY / ALL_PROXY → Windows system proxy → direct.
  • Empty string "" means explicitly direct (override-and-disable).
  • Auto mode honors NO_PROXY and the Windows proxy bypass list for environment/system fallbacks. A proxy pointing back to bili's own local port is ignored or rejected to prevent a loop.
  • HTTP and HTTPS proxy origins are supported. SOCKS5 (socks5/socks5h) is not supported: an explicit BILI_UPSTREAM_PROXY / config proxy with such a scheme fails startup with an actionable error, while env/system proxies (HTTPS_PROXY, …) with such a scheme are ignored with a one-time warning (traffic then falls through to direct). For Clash/mihomo, point bili at the same mixed port over http:// (e.g. http://127.0.0.1:7890).
  • Both outbound paths are covered: /bili/ path-mode (fetch) AND MITM CONNECT tunnels (the proxy's connection to the real upstream goes through the HTTP CONNECT proxy).
  • The auto-updater's own egress (npm registry check + tarball download) uses the same decision for its hosts, so bili update and auto-update work on hosts where npm is only reachable through the proxy (#609).

Env override: BILI_UPSTREAM_PROXY=http://127.0.0.1:20172 (higher priority than the config file). On Windows, common Clash/Mihomo static system proxies are discovered automatically; the Web UI shows the effective source and any PAC URL detected in Internet Settings.

MITM vs /bili/ — distinguishing the key scheme. A login client (ZCode via MITM) and an API-key client can both hit the same host (open.bigmodel.cn). To let their config differ, MITM traffic uses a mitm:// scheme in the lookup key while /bili/ traffic uses the real https://:

ClientLookup key example
ZCode (MITM, login)mitm://open.bigmodel.cn
API-key client (/bili/)https://open.bigmodel.cn/api/anthropic

So you can give ZCode its own proxy without affecting API-key clients:

{
  "providers": {
    "mitm://open.bigmodel.cn":            { "proxy": "http://127.0.0.1:20173" },
    "https://open.bigmodel.cn/api/anthropic": { "proxy": "http://127.0.0.1:20172" }
  }
}

Wire-compat role rewrite (compat.roles)

Some upstreams reject the developer role newer codex clients send on the Responses API (400 Invalid role: developer). compat.roles maps roles to what the upstream accepts — applied at the forward boundary to the final openai/responses body (client-sent roles and bili's own injected prompt alike), global or per-provider, default off = byte-for-byte:

{
  "compat": { "roles": { "developer": "system" } },
  "providers": {
    "https://picky.example.com": { "compat": { "roles": { "developer": "user" } } }
  }
}

No configuration needed for the common case. When an upstream answers a request with 400 Invalid role: …, bili auto-rewrites the offending role to system, retries the request once, and — if the retry succeeds — remembers the mapping for that session only (nothing is written to your config). Later requests in the session skip the 400 round-trip. The log line printed when the auto-fix fires includes a copy-paste per-provider snippet if you want the mapping permanently.