rasyidmmz/dsh-paper-search ↗★ 0
dsh-paper-search
集成16个学术源的文献检索工具 适合需要进行学术文献检索、论文状态查询的科研人员。
安装
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:rasyidmmz/dsh-paper-search说明文档
阅读完整 README ↗dsh-paper-search
Literature search for DeepSeek Harness: sixteen sources behind one tool, all native HTTP — no external CLI, no Python package. Fourteen of them need no credential at all.
Registered as the paper-search skill and two model tools (paper_search,
paper_status).
Sources
International (13) — official JSON APIs, except IACR which is parsed HTML:
| id | Source | Notes |
|---|---|---|
crossref | Crossref | DOI metadata across publishers, broadest coverage |
openalex | OpenAlex | metadata + abstracts; an email raises the rate limit |
doaj | DOAJ | open-access journals — includes many Indonesian journals |
europepmc | Europe PMC | biomedicine, full text for OA records |
pubmed | PubMed | biomedicine; two calls per search (esearch + esummary) |
pmc | PubMed Central | open-access full text; two calls per search |
arxiv | arXiv | preprints; the plugin self-paces to 1 request / 3 s per arXiv's TOU |
semantic | Semantic Scholar | a free key raises the rate limit |
core | CORE | open-access repository aggregator — CORE_API_KEY required |
zenodo | Zenodo | general repository: data, software, preprints |
hal | HAL | French open-access repository |
iacr | IACR ePrint | cryptology preprints |
unpaywall | Unpaywall | DOI input, not keyword — resolves a DOI to its open-access copy; email required |
Indonesian (3) — parsed from public pages:
| id | Source | Notes |
|---|---|---|
garuda | Garuda (Kemdiktisaintek) | Indonesian journal articles |
sinta | Sinta (Kemdiktisaintek) | accredited journals, not articles |
ios | IOS OneSearch (Perpusnas) | Indonesian repository aggregator |
sources: "all" and the DOI-input source
unpaywall answers only when the query is a DOI. So when you search with
keywords, it is skipped rather than run and reported as an empty result — and
the output says so:
CATATAN: unpaywall dilewati karena query ini bukan DOI — sumber itu memang
mencari dengan DOI, bukan kata kunci.
Pass a DOI and it runs. Pass sources: "unpaywall" with keywords and it runs
too, returning nothing — because that is genuinely how it behaves, and the
output says that rather than pretending the source is broken.
Deliberately excluded, with the reason
Every exclusion was measured on 2026-09-28. An always-empty source is worse than an absent one, so these are dropped rather than registered to look larger:
| Source | Reason |
|---|---|
moraref | landing page only, no parseable result markup (3,964 characters) |
dblp | API is behind bot protection — returns "Making sure you're not a bot!", not JSON |
openaire | test request timed out at 30 s and again at 40 s |
biorxiv, medrxiv | their API is date-range based, not keyword search |
base | OAI-PMH requires institutional IP registration (Access denied for IP address …) |
citeseerx | returns HTTP 404, then hangs until the request is killed |
ssrn | rejects ordinary requests with HTTP 403 |
google_scholar | bot detection active; a proxy is required |
Install
# from npm
dsh plugin --profile
add dsh-paper-search
# or straight from the repository
dsh plugin --profile
add github:rasyidmmz/dsh-paper-search
The package ships a dsh.bundle.patch, so DSH installs its loader row
automatically. Do not also write an id: paper-search row by hand — the
loader throws duplicate loader entry id rather than warning.
Restart DSH after installing.
Use
paper_search({ query: "machine learning" })
paper_search({ query: "pendidikan karakter", sources: "indonesia" })
paper_search({ query: "CRISPR", sources: "pubmed,europepmc,doaj", max_results: 10 })
paper_search({ query: "digital literacy", open_access_only: true, year_from: 2020 })
paper_status() # which sources answer right now
paper_status({ probe: false }) # configuration only, no network
For Indonesian topics, use Indonesian keywords: "pendidikan karakter"
returns real Indonesian journals, while the English equivalent returns
international results that do not match the intent.
Verification status
Stated precisely, because "tested" on its own is not a useful claim:
Verified — the connectors, against live APIs and pages. node test/uji-sumber.mjs
runs all sixteen against the real services: 8/8 assertions. On the last run 15 of
16 answered. The one that did not was PubMed Central returning a transient
HTTP 500 error forwarding request — and it was reported as a failure, which is
the behaviour this package exists to provide.
Verified — the wiring, against a mock Cordis context. node test/uji-plugin.mjs,
29/29 assertions: both tools register, the skill provider reports rank 450,
the frontmatter parses within the 500-character catalogue limit, and both tools
reach the network end to end.
Verified — loading the bundle inside a live harness. node test/uji-di-dsh.mjs
copies the headless profile to a throwaway test profile, inserts this plugin
plus a witness plugin, boots a real DSH, and checks from inside that both
tools are registered, the skill provider is registered, the skill appears in the
catalogue, and paper_search is actually called and returns real papers.
22 assertions, 0 failures.
That last test earned its place: it caught a fatal bug the mock never could.
DSH requires every tool to declare output: { schema, render }, and 0.1.0 did
not — so registration threw, and because the throw happened inside apply(),
the whole plugin failed to load. 0.1.0 is unusable; use 0.1.1.
What makes the reporting honest
Every result separates sources that failed from sources that answered:
v Crossref 3 results (2316 ms)
x Semantic Scholar FAILED — HTTP 429 — {"message": "Too Many Requests...
v Garuda 3 results (1428 ms)
A source that fails is never allowed to masquerade as "no results". This is not
a detail: a widely-used tool in this space returns an empty list for every
non-200 response with no error recorded at all, so a rate-limited request reads
as "this literature does not exist". Here, the failure and its HTTP status are
printed, and paper_status exists to prove liveness on demand.
Credentials (all optional)
No key is bundled. Read order:
- plugin config in your DSH profile,
- environment variables
PAPER_SEARCH_MCP_, then ``, - the file
~/.config/paper-search-mcp/.env(belonging to the upstream CLI).
| Name | Purpose |
|---|---|
UNPAYWALL_EMAIL | any email — used as the Crossref/OpenAlex "polite pool" contact |
OPENALEX_EMAIL | same, OpenAlex only |
SEMANTIC_SCHOLAR_API_KEY | raises Semantic Scholar rate limits |
DOAJ_API_KEY | raises DOAJ's hourly limit |
CORE_API_KEY | required for the core source — without it CORE answers HTTP 429 with an empty body |
Fourteen of the sixteen sources work with no credential at all; keys raise rate limits and stability. Two are exceptions, and each is reported as a clear, actionable failure rather than a confusing 429:
coreneedsCORE_API_KEY.unpaywallneeds an email (any real address) — it also needs a DOI as input.
What this package does not do
- It does not download or redistribute full-text PDFs.
- It does not bypass paywalls or access controls.
- It does not call any service other than the ten listed above.
- It does not require or install
paper-search-mcp. If that CLI happens to be installed, you can use it for its additional sources (CORE, Zenodo, HAL, SSRN, Unpaywall, OpenAIRE, CiteSeerX, BASE) — this plugin never depends on it.
Licence and attribution
MIT — see LICENSE. Attribution and source terms are in NOTICE.md.