Evaluation¶
Harness Evals¶
The eval suite (evals/) measures whether the harness is agent-legible: can a cold-start agent answer basic operational questions correctly from the project's AGENTS.md and docs?
How it works¶
Each task is a row in evals/tasks.tsv:
| column | description |
|---|---|
id |
short identifier |
prompt |
question asked of the agent via opencode run |
pattern |
grep regex the answer must match to pass |
The runner (evals/run.sh) invokes opencode run with a cheap OpenRouter model for each task and reports pass/fail. Running from the project root means the agent has AGENTS.md and the doc tree as context — the same cold-start position a new contributor agent would have.
Running locally¶
Override the model:
Save a timestamped result file to a directory:
A task that does not match its pattern is recorded as FAIL and does not make
the runner exit non-zero. The runner exits non-zero only when the harness
itself errors (opencode fails or times out).
CI¶
The harness-evals.yml workflow runs on manual dispatch. It uses
OPEN_ROUTER_API_KEY from repository secrets. After each run,
evals/publish-results.sh commits the timestamped result .tsv file to the
unprotected eval-results
branch, one file per run, for longitudinal tracking of the pass rate. Results
are never committed to main, which is protected.
Adding tasks¶
Add a tab-separated row to evals/tasks.tsv. Keep prompts as a cold-start agent would ask them (no project-specific jargon assumed). Patterns are extended regex (grep -E).
Result files are on the eval-results branch, not in this directory.