uson1x/dsh-plugin-llm-verifier0

dsh-plugin-llm-verifier

LLM-as-a-Verifier for DeepSeek Harness (dsh): continuous reward signals via select / compare / track, using Monte Carlo repeated grading over decomposed criteria

AI Analysis

核心用途是利用大模型作为评审器,对 Agent 的输出进行多维度、高精度的自动打分与对比。适合需要对模型生成质量进行自动化评估和对齐训练的开发者。

Package
dsh-plugin-llm-verifier
Version
0.1.0
License
MIT
Last updated
Aug 18, 2026

Install

This plugin has no verified bundle, or compatibility checks failed. Read the repository notes first. Read the full README ↗

Configuration

KeyDefaultMeaning
provider— (required)Registered ctx.llm provider route to grade with
model— (required)Model id on that route
granularity20Integer score scale 1..G
repetitions3Grading passes per criterion (K)
temperature1Sampling temperature for the Monte Carlo estimate
criteriaspecification / output / errorsArray of { name, description } sub-criteria
maxOutputTokens2048Output cap per grading call
timeoutMs120000Deadline per grading call
concurrency4Parallel grading calls
tieMargin0.02compare margin below which the verdict is tie

Cost note: verify_select makes N × C × K model calls (default 9 per candidate); verify_track makes steps × K.