jaco-tech/dsh-web-fetch-crw ↗★ 0
@jaco-tech/dsh-web-fetch-crw
crw (Firecrawl-compatible) fetch provider for the DeepSeek Harness web capability seam (ctx.web). No auth, no API key.
AI Analysis
核心用途是为 DSH 的网页获取工具提供抓取服务,将网页内容转换为 Markdown。适合需要 AI 抓取和分析网页内容的用户。需要自行托管 crw 实例。
Install
This plugin has no verified bundle, or compatibility checks failed. Read the repository notes first. Read the full README ↗
README
Read the full README ↗dsh-web-fetch-crw
crw (Firecrawl-compatible) fetch provider plugin for DeepSeek Harness.
Registers a WebFetchProvider into ctx.web so that the existing web_fetch model tool retrieves page content through your crw instance — which converts HTML to clean markdown. No auth, no API key.
crw is a self-hosted web scraping and crawling service that implements the Firecrawl API. It provides a drop-in replacement for the Firecrawl /v1/scrape endpoint.
Status
⚠️ Not yet published to npm. The CI/CD workflow builds on every push and publishes automatically when a v* tag is pushed. Once the first release is cut:
npm install @jaco-tech/dsh-web-fetch-crw
Configure
Add to your DSH profile's cordis.patch.yml:
- insert:
- id: web-fetch-crw
name: '@jaco-tech/dsh-web-fetch-crw'
config:
baseURL: 'http://your-crw-instance:3000'
Optionally pin ctx.web to use this provider explicitly:
- replace:
- id: web
config:
fetchProvider: crw
How it works
This is a Cordis plugin that implements the WebFetchProvider interface from @deepseek-ai/dsh-web:
export const name = "web-fetch-crw"
export const inject = ["web"]
export function apply(ctx, config) {
ctx.web.registerFetchProvider({
id: "crw",
available() { /* URL.canParse(config.baseURL) */ },
async fetch(request, signal) {
// POST /v1/scrape { url, formats: ["markdown"] }
// maps response to WebFetchResult (body as markdown text)
}
})
}
The existing @deepseek-ai/dsh-tool-web consumer calls ctx.web.fetch() — your crw instance serves the content automatically.
Provider interface
| Method | Behavior |
|---|---|
available() | Returns true when baseURL is a valid URL |
fetch(request, signal?) | POST /v1/scrape with { url, formats: ["markdown"] }. Returns the scraped markdown as kind: "text" body (no double-conversion through turndown). |
Upstream
This plugin communicates with crw via its Firecrawl-compatible API (POST /v1/scrape). crw is a self-hosted web scraping and crawling service that renders pages (via Lightpanda or Chromium) and produces clean markdown. You need a running crw instance to use this plugin.
Release process
# Tag and push — the CI workflow builds and publishes to npm automatically
git tag v0.1.0
git push origin v0.1.0
Requires the NPM_TOKEN secret to be set in the repository.
Development
npm install
npm run build
License
MIT