Skip to content

Evaluation

Harness Evals

The eval suite (evals/) measures whether the harness is agent-legible: can a cold-start agent answer basic operational questions correctly from the project's AGENTS.md and docs?

How it works

Each task is a row in evals/tasks.tsv:

column description
id short identifier
prompt question asked of the agent via opencode run
pattern grep regex the answer must match to pass

The runner (evals/run.sh) invokes opencode run with a cheap OpenRouter model for each task and reports pass/fail. Running from the project root means the agent has AGENTS.md and the doc tree as context — the same cold-start position a new contributor agent would have.

Running locally

OPENROUTER_API_KEY=... bash evals/run.sh

Override the model:

MODEL=openrouter/xiaomi/mimo-v2.5 bash evals/run.sh

Save a timestamped result file to a directory:

RESULTS_DIR=/tmp/eval-results bash evals/run.sh

A task that does not match its pattern is recorded as FAIL and does not make the runner exit non-zero. The runner exits non-zero only when the harness itself errors (opencode fails or times out).

CI

The harness-evals.yml workflow runs on manual dispatch. It uses OPEN_ROUTER_API_KEY from repository secrets. After each run, evals/publish-results.sh commits the timestamped result .tsv file to the unprotected eval-results branch, one file per run, for longitudinal tracking of the pass rate. Results are never committed to main, which is protected.

Adding tasks

Add a tab-separated row to evals/tasks.tsv. Keep prompts as a cold-start agent would ask them (no project-specific jargon assumed). Patterns are extended regex (grep -E).

Result files are on the eval-results branch, not in this directory.