v0.5.0 · released 2026-09-30
Typed evaluation for OpenCode skills and AI-agent workflows. Run a skill with OpenCode, or grade content-addressed evidence from another runner without starting a model. Results preserve behavior status, product dimensions, evidence availability, and provenance instead of flattening them into one score.
- Capability, regression, and product evaluations. Each case declares what it measures and whether it compares against a baseline.
- Typed outcomes.
PASS,FAIL,ERROR,NEEDS_REVIEW, andINDETERMINATEstay distinct. Missing evidence is never a pass or a measured zero. - Provider-neutral grading.
gradevalidates a schema-1 manifest binding the case, workdir tree, transcript, artifacts, environment, and provenance. It does not execute a model or load arbitrary evaluator plugins. - Independent product dimensions. Numeric measurements stay grouped by their declared dimension; eval-harness does not compute a global product-quality score.
- Auditable comparisons and reports. Baselines, typed A/B results, attribution, token/cost/duration coverage, JUnit XML, and SARIF help explain what changed and what evidence is available.
run is the OpenCode execution adapter. A LangGraph runner example demonstrates the runner contract; grade accepts prepared evidence without coupling the grader to that runner.
The package is distributed from GitHub; it is not published to the npm registry. Install the checked-out source and link its CLI:
git clone https://github.com/nano-step/eval-harness.git
cd eval-harness
npm link
eval-harness --versionNode.js 18+ is required for the CLI link. Running npm link creates a local symlink to this checkout; it does not publish a package.
Set OPENCODE_SKILLS_ROOT to the directory containing <skill-name>/evals/cases/*.yaml, then run:
export OPENCODE_SKILLS_ROOT="/path/to/skills-root"
eval-harness run --skill=my-skill --dry-run
eval-harness run --skill=my-skillUse --dry-run to check discovery and preflight without spawning an evaluation. After reviewing a passing run, record a baseline and enable strict gates:
eval-harness baseline --skill=my-skill
eval-harness run --skill=my-skill --strictProvider-neutral grading uses a deterministic evidence bundle and manifest:
eval-harness grade --manifest=grading-manifest.json --strictStrict grading exits 13 for malformed or tampered evidence, 14 for FAIL, 15 for pending human review, and 16 for indeterminate evidence. The grader never starts a model.
| Type | Use | Baseline comparison |
|---|---|---|
regression |
Preserve an existing behavior contract | Enabled by default |
capability |
Verify declared behavior or coverage | Disabled by default |
product |
Report named product measurements | Disabled by default |
Shipped checks: shell, jq_path_contains, file_exists, output_contains, output_not_contains, llm_judge, metric_score, trajectory, and human_review. Required checks determine the typed result; optional checks report evidence without gating. Product dimensions remain separate.
For prose checks, llm_judge requires ANTHROPIC_API_KEY; abstentions or unresolved votes remain unavailable rather than becoming a synthetic pass. Deterministic checks and grade do not require a model.
- Repeated trials report pass@k, pass^k, and Wilson bounds for the same case/configuration. The IID assumption is explicit, not measured.
- Token, cost, and duration fields carry measured/partial/unavailable coverage. Unknown cost remains
null; withEVAL_BUDGET_USDenabled, unmeasured ledger spend blocks later gated runs until reconciled. - Attribution reports changed evidence and environment fields. Hash co-occurrence is not proof that a change caused an outcome.
eval-harness abcompares two skills and can warn on cost increases; its cost threshold is warning-only.eval-harness metaevalruns an offline stub corpus and reports harness validity separately from skill results.
v0.5.0 includes a documented 0.x breaking change. By default, kind: shell accepts a constrained, shell-free jq/printf/wc -l language; jq input files must resolve inside the case workdir. Other shell commands must migrate to a typed grader or opt in with unsafe_shell: true (or EVAL_ALLOW_UNSAFE_SHELL=1) only when the case YAML is trusted. The opt-in runs with the harness user's permissions; the workdir is not an OS sandbox.
Run the deterministic offline suite with:
npm testThe suite uses fixtures and stub runners; it does not measure live model quality, latency, or cost. Optional exports are available with --report=junit:<path> and --report=sarif:<path>.
- Documentation index
- v0.5.0 design and migration contract
- Runner contract and LangGraph example
- Janus benchmark: offline evidence and the no-go decision on a native adapter
- Security policy, versioning policy, and known issues
- Changelog and contributing guide
eval-harness measures declared behavior and product evidence. It does not review skill-design quality, infer causality from hashes, or replace application unit tests. JUnit/SARIF reports are local outputs; the project has no hosted results service. See the Janus benchmark for why its experimental evaluator is not currently used as a native execution adapter.
MIT · nano-step