@zseven-w/dsh-qa
QA orchestrator plugin for DeepSeek Harness: an agent explores your app and leaves evidence, then the explored path replays deterministically. Drives the browser, macOS desktop, iOS and Android drivers. 适合测试人员,用于驱动浏览器、桌面或移动端进行自动化应用测试。
Install
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:ZSeven-W/dsh-qaREADME
Read the full README ↗DSH QA
Explore real apps, capture evidence, and turn verified actions into repeatable QA scenarios.
Agent-Led Exploration • Evidence-Backed Assertions • Deterministic Replay • Browser, Desktop, iOS & Android
Package: @zseven-w/dsh-qa · Version: 0.1.0-rc.2 · Prerelease
Quick start · Safety and limitations · Known gaps · Development · Documentation

Actual output from the shipped browser example: two steps and final assertions passed. This is a typeset view of the unedited Markdown report, not a built-in dashboard or a claim of four-platform coverage.
Why DSH QA
QA orchestrator plugin for DeepSeek Harness: an agent explores your App like a real user (with evidence-backed findings), then the explored path is exported as a deterministic Replay scenario that runs on every release.
- Explore mode — the agent drives the app through a tool surface
(
qa_session_start/qa_observe/qa_act/qa_assert/qa_evidence/qa_record_export/qa_replay_run/qa_session_stop) and documents what it sees. - Replay mode — declarative
QaScenariofiles (lossless JSON) executed deterministically with per-step re-observe assertions, producing redacted JSON / Markdown / JSONL reports. - Drivers —
@zseven-w/dsh-browser(BU, contract v9),@zseven-w/dsh-computer(CU, contract v5),@zseven-w/dsh-ios, and@zseven-w/dsh-android. Driver safety semantics are inherited, never loosened:EXTERNAL_COMMIT_TARGETrefused, secure fields permanently refused, approval gates passed through,unknownreceipts require re-observation. Mobile sessions require an explicit device id and never fall back to a default device. Mobile text input is conditional, never blanket-implemented: iOSfill/typeuse the dsh-ios native element-boundfillTarget/typeTargetprimitives when the live driver exposes them, with missing methods or missing native identifiers left explicitly unavailable and no raw global-type fallback; Androidtyperemains append-faithful after real focus verification and AndroidfillremainsFILL_PRIMITIVE_UNAVAILABLE.
QA coordinates four independent drivers; it does not replace them or grant broader access. A passing fixture, unit suite, or device acceptance run is evidence for that tested scope, not a claim that every app is supported.
Quick start
Requires Node.js 24.11.0 or later. QA orchestrates drivers it does not
contain: it loads each one lazily, by package name, only when a session asks
for that platform. Installing @zseven-w/dsh-qa alone gives you no
drivers — qa_session_start will report the driver as not installed. Install
the ones you need alongside it:
npm install @zseven-w/dsh-qa # the orchestrator
npm install @zseven-w/dsh-browser # browser sessions
npm install @zseven-w/dsh-computer # macOS desktop sessions
npm install @zseven-w/dsh-ios # iOS device/simulator sessions
npm install @zseven-w/dsh-android # Android device/emulator sessions
| Driver | Package | Responsibility | Extra prerequisites |
|---|---|---|---|
| Browser (BU) | @zseven-w/dsh-browser | Browser observation and interaction | An installed Chrome / Edge / Chromium; the driver discovers one and never downloads it |
| Computer (CU) | @zseven-w/dsh-computer | Native desktop observation and interaction | macOS; a locally built + granted Helper (Accessibility + Screen Recording) — see Known gaps |
| iOS | @zseven-w/dsh-ios | Explicit-device mobile sessions | macOS + Xcode; an explicit device id |
| Android | @zseven-w/dsh-android | Explicit-device mobile sessions | adb; an explicit device id |
Verify the install by replaying the shipped browser example, which needs
nothing but dsh-qa, dsh-browser, and a local browser:
node node_modules/@zseven-w/dsh-qa/scripts/run-example.mjs
It serves the packaged web fixture on an ephemeral loopback port, replays a
two-step scenario headlessly, and writes JSON / Markdown / JSONL reports. A
working install prints [example] status: pass. See
Running the shipped example for the details.
For DSH-host usage, install or update DSH with:
npm install -g @deepseek-ai/dsh@latest
Installing DSH does not activate this plugin. Its host entry is declared in
cordis.patch.yml; the standalone stdio MCP entry is
src/server.mjs, also exposed by npm run mcp and
.mcp.json. Configure the chosen host to load the plugin/server
and provide the required drivers before starting a session. Follow the
Explore playbook for the observe → act → assert
→ evidence → export workflow.
Safety and limitations
- No false green:
unknownaction receipts need fresh proof. Runs reportpass,inconclusive, orfail; missing coverage is not absence. - No authority escalation: driver approvals remain in force. Secure fields
and
EXTERNAL_COMMIT_TARGETare refused; mobile sessions require an explicit device id, with no default-device fallback. - Replay needs durable targets: coordinates, ephemeral references, and ambiguous selectors are not promoted into durable scenarios. An action with no settled, provable outcome is excluded from export.
- Vision is advisory: visual assertions assist triage but do not determine the run status. Model narration is not observed fact.
- Mobile support is conditional: iOS text input requires live native
element-bound primitives and identifiers. Android
typeverifies real focus; Androidfillremains unavailable. - Evidence needs care: structured reports use fail-closed redaction, but screenshots can still contain private content. Use synthetic test data and inspect artifacts before sharing. Login-state injection requires explicit owner authorization for exact origins; see login state.
Known gaps
This is a developer release. The main path — Explore → evidence → Export → Replay on browser and desktop — is exercised by the suite on every change, but these are open, and knowing them is part of using the package honestly.
- The Computer helper is built by you, by design. It ships ad-hoc signed
with no TeamIdentifier and no stapled ticket: you build it from the
dsh-computercheckout and grant it Accessibility + Screen Recording yourself. This is a settled decision, not a pending task — there is no Developer ID notarization planned, so do not wait for an "install and go" desktop build. The browser, iOS and Android drivers are unaffected. - Absence is often unproven rather than proven, and that shows up as
inconclusive.node-absentpasses only when the deciding view is complete AND the browser driver's bounded closed-shadow-root probe verified its coverage. When that probe finds a closed root, exceeds its node budget, or cannot open a CDP session, the assertion carriesCOVERAGE_UNVERIFIEDorINCONCLUSIVE_TRUNCATEDand the run isinconclusive— neverpass. On shadow-DOM-heavy or very large pages, expect that instead of a green absence.failis reserved for a real defect: a node that IS observed disprovesnode-absentand fails normally, whatever the coverage state. - Deep targets on large pages stay inconclusive. Beyond the driver's
100-node observation window, a scoped scroll proof can only reach
INCONCLUSIVE_SCOPE, neverpass. This is honest, not broken — but it means deep flows on big pages do not produce a green gate today. - Visual assertions are advisory and the live vision path is unverified
here. Vision never changes a Replay's pass/fail by design. The seam to a
real host vision service (
ctx.llm/attachments) is covered only by a fake in the suite; it has not been run against a live vision model. - Mobile drivers carry no contract version. The browser and computer
drivers publish a contract version and QA now REFUSES to load one that is not
the version this build supports (browser 9, computer 5) — older and newer
alike, because a contract change can alter what an observation means rather
than merely adding to it.
dsh-ios/dsh-androidexport no version from/driver, so QA still loads them structurally and a drift there is caught after the fact rather than at load.
Replay verifies five kinds of semantic assertion. A pass means those
assertions held on fresh observations — not that the app is correct, not that
the screen looks right, and not that coverage was complete.
Implementation reference
Session semantics, replay, observation coverage, and host integration
Session core
src/session/ implements the QA loop observe -> act -> re-observe -> evaluate -> evidence -> cleanup on top of a driver adapter interface, with these hard rules:
- an
unknownaction receipt is never treated as success — the outcome is decided only by a fresh observation; rejected/failedreceipts propagate as step failures with the receipt attached as evidence;- every session cleans up (
driver.stop) even on failure.
src/adapters/browser.ts adapts @zseven-w/dsh-browser (declared as a
link:../dsh-browser dev-only devDependencies linkage, never a runtime
dependency and never vendored) to that interface
without weakening any driver safety semantics. fixtures/web/index.html is a
self-contained loopback fixture that reproduces the 2026-08-25 acceptance flow.
The native (Computer-driver) fixture fixtures/native/ is repository-only by
decision (QA-BL-042, 2026-09-05): it is a signed macOS app bundle built from
main.swift and is not in the published package. Build it from a checkout with
node fixtures/native/build-fixture.mjs, launch
fixtures/native/build/DshQaFixture.app, and stop it with
pkill -x dsh-qa-fixture. It is the only isolated target for demonstrating the
Computer driver's permanent secure-field rejection ("Secure password",
fixture.securePassword) and is exercised by test/computer-integration.test.mjs.
Replay
Declarative QaScenario files (lossless JSON, { meta, target, steps[], assertions[] }) are executed deterministically by src/replay/runner.ts on
top of the session core. Every step carries the act plus the assertion that
must hold on the FRESH observation after it; an unknown receipt is never a
pass — only the re-observation decides. The fail-closed loader
(src/replay/loader.ts) rejects unknown fields, malformed steps, missing
fields, and non-lossless values without echoing value bytes. Reporters emit
redacted report.json / report.md / append-only report.jsonl through
the fail-closed v2 engine in src/redaction.
See scenarios/examples/ and qa_assert / qa_replay_run.
Explore → Replay
src/explore/ wraps the existing driver adapter as a passive recorder. It records
ordered observations, actions, receipts, and evidence references after applying the
fail-closed redaction projection; the session core itself is unchanged. Ephemeral
driver refs are replaced by per-session correlation aliases and never enter the
trajectory.
qa_record_export accepts an output_path under the current workspace or temporary
directory, synthesizes every step assertion from the SETTLED fresh observation
after that action, writes a scenario, and reads the exact bytes back through the
existing fail-closed loader. Rejected/failed actions, actions without a fresh
observation, and actions whose view never settled are returned in
excludedActions, never silently promoted to steps.
Selector durability is intentionally strict. Browser targets require a unique
non-empty role plus accessible name; computer targets use the durable
Accessibility identifier; mobile targets prefer the stable Android resourceId /
iOS AXUniqueId identifier. Indices, coordinates, observation ids, generated ids,
duplicate names, unnamed nodes, and ephemeral refs are refused instead of
guessed. An unknown receipt additionally needs an observable semantic delta or
URL change; the old target merely remaining present is not proof.
The Explore methodology ships as two artifacts: a human-readable copy at
skills/qa-explore/SKILL.md (included in the package files), and the playbook
the plugin actually registers through the optional skill service — the
QA_SKILL_CONTENT template literal in src/skill.ts, registered under the
name qa-orchestration via ctx.inject(['skills'], …).
Bounded settle (asynchronous UIs)
Real UIs are asynchronous, so a single proof observation taken immediately after
an action is a race: it can miss the outcome that has not rendered yet, or catch
unrelated late hydration churn and mistake it for the outcome. src/session/settle.ts
replaces that single-shot read with a bounded settle: observe until the semantic
projection (refs and other session-local identity excluded) has held still for
QA_SETTLE_QUIET_MS, bounded by QA_SETTLE_BUDGET_MS, and — right after an
action — never conclude from silence alone.
The SAME policy is applied by the session core to Explore's proof observations
and Replay's verification observations, so the two sides can never judge
different views of the same page. A view that never settles is honestly
unprovable: export excludes the step (ASSERTION_NOT_PROVABLE) and replay fails
the step. Nothing is widened to make an unstable page pass. Assertion synthesis
prefers evidence on or near the action target and records the weakness in the
step intent when only a distant delta exists. A fill's own value echo on its
target is masked from the change decision (it is expected, not evidence of a
downstream outcome), and the exporter proves the fill with a node-value
assertion on the target whenever the settled view shows it carrying the typed
text.
Configure it with new QaToolHost({ settle: ... }), runScenario(..., { settle: ... }),
or the DSHPLUGIN_QA_SETTLE_BUDGET_MS / DSHPLUGIN_QA_SETTLE_QUIET_MS / DSHPLUGIN_QA_SETTLE_INTERVAL_MS
environment overrides. Full rationale, both reproduced real-world failure modes,
and the regression fixtures: docs/SETTLE.md.
A truncated view is incomplete, not empty
Observations are budget-limited and carry truncated. A node beyond the budget
still exists, so an absence can never be proven from a truncated view:
node-absent fails closed there instead of reporting a silent false green, and
any outcome that would rest on an incomplete view re-observes ONCE at
QA_ESCALATED_NODE_BUDGET before concluding. A claim that is still unprovable
against the fuller view carries completeness.reason: "INCONCLUSIVE_TRUNCATED"
— distinct from an ordinary failure — while a node that WAS returned remains
sound evidence of presence. Whenever truncation touched a decision, the
completeness block (budget, truncation state, escalation) is written into
report.json, report.jsonl and a "view completeness" line in report.md.
The