rasyidmmz/dsh-paper-search ↗★ 0

dsh-paper-search

集成16个学术源的文献检索工具 适合需要进行学术文献检索、论文状态查询的科研人员。

包名
dsh-paper-search
兼容性
待验证
版本
0.1.1
许可证
MIT
最近更新
2026年9月28日

安装

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:rasyidmmz/dsh-paper-search

dsh-paper-search

Literature search for DeepSeek Harness: sixteen sources behind one tool, all native HTTP — no external CLI, no Python package. Fourteen of them need no credential at all.

Registered as the paper-search skill and two model tools (paper_search, paper_status).

Sources

International (13) — official JSON APIs, except IACR which is parsed HTML:

idSourceNotes
crossrefCrossrefDOI metadata across publishers, broadest coverage
openalexOpenAlexmetadata + abstracts; an email raises the rate limit
doajDOAJopen-access journals — includes many Indonesian journals
europepmcEurope PMCbiomedicine, full text for OA records
pubmedPubMedbiomedicine; two calls per search (esearch + esummary)
pmcPubMed Centralopen-access full text; two calls per search
arxivarXivpreprints; the plugin self-paces to 1 request / 3 s per arXiv's TOU
semanticSemantic Scholara free key raises the rate limit
coreCOREopen-access repository aggregator — CORE_API_KEY required
zenodoZenodogeneral repository: data, software, preprints
halHALFrench open-access repository
iacrIACR ePrintcryptology preprints
unpaywallUnpaywallDOI input, not keyword — resolves a DOI to its open-access copy; email required

Indonesian (3) — parsed from public pages:

idSourceNotes
garudaGaruda (Kemdiktisaintek)Indonesian journal articles
sintaSinta (Kemdiktisaintek)accredited journals, not articles
iosIOS OneSearch (Perpusnas)Indonesian repository aggregator

sources: "all" and the DOI-input source

unpaywall answers only when the query is a DOI. So when you search with keywords, it is skipped rather than run and reported as an empty result — and the output says so:

CATATAN: unpaywall dilewati karena query ini bukan DOI — sumber itu memang
mencari dengan DOI, bukan kata kunci.

Pass a DOI and it runs. Pass sources: "unpaywall" with keywords and it runs too, returning nothing — because that is genuinely how it behaves, and the output says that rather than pretending the source is broken.

Deliberately excluded, with the reason

Every exclusion was measured on 2026-09-28. An always-empty source is worse than an absent one, so these are dropped rather than registered to look larger:

SourceReason
morareflanding page only, no parseable result markup (3,964 characters)
dblpAPI is behind bot protection — returns "Making sure you're not a bot!", not JSON
openairetest request timed out at 30 s and again at 40 s
biorxiv, medrxivtheir API is date-range based, not keyword search
baseOAI-PMH requires institutional IP registration (Access denied for IP address …)
citeseerxreturns HTTP 404, then hangs until the request is killed
ssrnrejects ordinary requests with HTTP 403
google_scholarbot detection active; a proxy is required

Install

# from npm
dsh plugin --profile 
 add dsh-paper-search

# or straight from the repository
dsh plugin --profile 
 add github:rasyidmmz/dsh-paper-search

The package ships a dsh.bundle.patch, so DSH installs its loader row automatically. Do not also write an id: paper-search row by hand — the loader throws duplicate loader entry id rather than warning.

Restart DSH after installing.

Use

paper_search({ query: "machine learning" })
paper_search({ query: "pendidikan karakter", sources: "indonesia" })
paper_search({ query: "CRISPR", sources: "pubmed,europepmc,doaj", max_results: 10 })
paper_search({ query: "digital literacy", open_access_only: true, year_from: 2020 })
paper_status()                    # which sources answer right now
paper_status({ probe: false })    # configuration only, no network

For Indonesian topics, use Indonesian keywords: "pendidikan karakter" returns real Indonesian journals, while the English equivalent returns international results that do not match the intent.

Verification status

Stated precisely, because "tested" on its own is not a useful claim:

Verified — the connectors, against live APIs and pages. node test/uji-sumber.mjs runs all sixteen against the real services: 8/8 assertions. On the last run 15 of 16 answered. The one that did not was PubMed Central returning a transient HTTP 500 error forwarding request — and it was reported as a failure, which is the behaviour this package exists to provide.

Verified — the wiring, against a mock Cordis context. node test/uji-plugin.mjs, 29/29 assertions: both tools register, the skill provider reports rank 450, the frontmatter parses within the 500-character catalogue limit, and both tools reach the network end to end.

Verified — loading the bundle inside a live harness. node test/uji-di-dsh.mjs copies the headless profile to a throwaway test profile, inserts this plugin plus a witness plugin, boots a real DSH, and checks from inside that both tools are registered, the skill provider is registered, the skill appears in the catalogue, and paper_search is actually called and returns real papers. 22 assertions, 0 failures.

That last test earned its place: it caught a fatal bug the mock never could. DSH requires every tool to declare output: { schema, render }, and 0.1.0 did not — so registration threw, and because the throw happened inside apply(), the whole plugin failed to load. 0.1.0 is unusable; use 0.1.1.

What makes the reporting honest

Every result separates sources that failed from sources that answered:

  v Crossref            3 results  (2316 ms)
  x Semantic Scholar    FAILED — HTTP 429 — {"message": "Too Many Requests...
  v Garuda              3 results  (1428 ms)

A source that fails is never allowed to masquerade as "no results". This is not a detail: a widely-used tool in this space returns an empty list for every non-200 response with no error recorded at all, so a rate-limited request reads as "this literature does not exist". Here, the failure and its HTTP status are printed, and paper_status exists to prove liveness on demand.

Credentials (all optional)

No key is bundled. Read order:

  1. plugin config in your DSH profile,
  2. environment variables PAPER_SEARCH_MCP_, then ``,
  3. the file ~/.config/paper-search-mcp/.env (belonging to the upstream CLI).
NamePurpose
UNPAYWALL_EMAILany email — used as the Crossref/OpenAlex "polite pool" contact
OPENALEX_EMAILsame, OpenAlex only
SEMANTIC_SCHOLAR_API_KEYraises Semantic Scholar rate limits
DOAJ_API_KEYraises DOAJ's hourly limit
CORE_API_KEYrequired for the core source — without it CORE answers HTTP 429 with an empty body

Fourteen of the sixteen sources work with no credential at all; keys raise rate limits and stability. Two are exceptions, and each is reported as a clear, actionable failure rather than a confusing 429:

  • core needs CORE_API_KEY.
  • unpaywall needs an email (any real address) — it also needs a DOI as input.

What this package does not do

  • It does not download or redistribute full-text PDFs.
  • It does not bypass paywalls or access controls.
  • It does not call any service other than the ten listed above.
  • It does not require or install paper-search-mcp. If that CLI happens to be installed, you can use it for its additional sources (CORE, Zenodo, HAL, SSRN, Unpaywall, OpenAIRE, CiteSeerX, BASE) — this plugin never depends on it.

Licence and attribution

MIT — see LICENSE. Attribution and source terms are in NOTICE.md.