One coding agent. Twelve historical tasks. Three inference routes. Sixty durable cells.
Read the Blueprint writeup · Open the interactive flight recorder
Reproducible coding-agent evaluations with isolated Git fixtures, held-out graders, route-level telemetry, spend controls, and publication audits.
This repository contains the tooling behind the published 60-run local-versus-cloud coding experiment. The interactive flight recorder exposes every task, route, repetition, patch outcome, timeout, provider recovery, and billed dollar.
The code is currently private while the public release surface is reviewed. The intended license is MIT.
The harness asks one coding agent to solve the same repository tasks through multiple model routes. Each route can point at a local Anthropic-compatible server or a hosted gateway. The runner preserves the route as a complete system: model weights, quantization, context limits, provider selection, caching, serving runtime, and failure recovery all remain part of the measured condition.
Each cell produces separate records for:
- the first patch and optional repair patch;
- held-out command results;
- patch-integrity violations;
- agent process duration and deadline termination;
- request count, token classes, cache traffic, and first-byte latency;
- upstream failures and provider fallback metadata;
- estimated list-price cost and exact gateway billing metadata when supplied;
- power-source state and a bounded machine snapshot.
Two outcome lenses stay separate throughout the system:
| Lens | Requirement |
|---|---|
| Held-out checks | Every required grader command passes after held-out tests are overlaid. |
| Strict acceptance | Held-out checks pass and the patch does not alter tracked tests or dependency lockfiles. |
That separation shows when an agent produced executable behavior while modifying protected files.
Every case in the published 60-cell experiment has a held-out executable grader. The exact inventory is:
| Grader source | Published coverage | Storage |
|---|---|---|
| Historical repository tests | 11 cases, 13 test files | hiddenFiles in each case.json points to an exact path at the pinned targetCommit. The harness reads that Git blob during grading. |
| Benchmark-authored oracle | 1 case, 1 test file | graders/portfolio-editorial-toc.test.tsx is tracked directly in this repository and referenced through benchmarkHiddenFiles. |
The public calibration example adds three inspectable benchmark-authored graders under graders/public-example/. They independently check the normalized output contract, private-field exclusion, and malformed-input behavior.
Both grader forms are copied into a separate workspace only after the coding agent exits. The agent fixture contains the base commit and never receives benchmark-authored or target-commit tests. The eight ctx cases require access to their private source repository to recover their nine historical test files. The generated public fixture exercises the complete preparation and grader-calibration path without those repositories.
| Property | Implementation |
|---|---|
| Future-history isolation | A fixture is exported from one base commit, reinitialized as a new repository, and created without remotes or later Git objects. |
| Hidden graders | Target tests or benchmark-owned oracle files are overlaid only after the agent process exits. |
| Grader calibration | Every case must reject the known-bad base and accept the historical reference patch before matrix execution. |
| Route parity | Every route receives the same prompt, agent binary, tool list, process limits, and grading commands. |
| Counterbalanced order | Route order rotates across cases and repetitions to reduce simple time-of-day ordering effects. |
| Bounded execution | First pass, repair pass, grader commands, first byte, stream idle, and process termination have explicit deadlines. Hosted runs have per-run request-admission thresholds and a total run-reservation cap. |
| Durable cells | A run directory contains the captured patch, grades, relay ledger, process summary, and final normalized record. Completed cells are skipped on restart. |
| Frozen inputs | A committed matrix lock binds the effective manifest, schedule, agent binary, engine files, cases, graders, and configured model/runtime artifacts. |
| Private raw traces | Agent event streams and relay ledgers live under ignored directories with mode 0600; reports carry normalized fields and hashes. |
| Independent review | A redacted, hash-pinned evidence packet can be sent to two isolated reviewer models using a fixed audit rubric. |
flowchart LR
M[Manifest + case files] --> P[Fixture preparation]
P --> V[Known-bad and reference validation]
V --> F[Matrix freeze]
F --> S[Counterbalanced scheduler]
S --> A[Claude Code process]
A --> R[Local relay]
R --> L[Local model route]
R --> G[Hosted gateway route]
A --> D[Captured Git patch]
D --> H[Hidden grader overlay]
H --> O[Strict + held-out outcomes]
R --> T[Usage, provider, latency, cost ledger]
O --> Q[Reconciliation + reports]
T --> Q
Q --> E[Redacted evidence packet]
E --> X[Independent methodology reviewers]
Version 0.1 drives Claude Code in non-interactive mode and routes its Anthropic Messages API traffic through the local relay. Model endpoints must accept that protocol. The relay can target:
- a local Anthropic-compatible HTTP server;
- Vercel AI Gateway;
- another hosted endpoint that accepts Anthropic Messages requests and the configured bearer credential.
The runner does not currently provide adapters for Codex CLI, Gemini CLI, OpenAI Responses API agents, or arbitrary agent processes. Those adapters belong behind the same run-record contract and are planned as separate modules.
- Node.js 24 or newer
- Git
tarlsoffor managed local-runtime lifecycle commands- Claude Code at an exact pinned path and version
- local clones containing every configured base and target commit
- an AI Gateway credential for hosted routes
- macOS for the included power-source gate and local-runtime example; set
environment.requireAcPowertofalseon other systems
Core tests, scheduling, grading helpers, report normalization, and relay tests run without model credentials.
git clone [email protected]:zackproser/coding-agent-eval-harness.git
cd coding-agent-eval-harness
npm ci
npm run example:validateThe example command creates a deterministic Git repository under .state/public-example/, checks its manifest, prepares an isolated base fixture, verifies that three held-out graders reject the known-bad commit, and verifies that they accept the reference commit. It does not invoke a coding agent, contact a model endpoint, or require credentials.
The root manifest preserves the exact published experiment configuration. Eight ctx cases point to a private source repository and four portfolio cases point to a public repository. The separate fixture under examples/public-fixture/ is the credential-free starting point.
To configure a measured run, copy the environment template and supply your agent, repository, runtime, model, and gateway paths:
cp .env.example .envRun the offline checks first:
npm test
npm run schedule
npm run doctordoctor checks environment substitutions, agent version, route configuration, source repositories, base and target commits, grader declarations, schedule uniqueness, and credential availability. It never sends a model request.
Prepare and calibrate one case:
npm run fixtures:prepare -- --case ctx-uuid-validation
npm run validate -- --case ctx-uuid-validationThen prepare and validate the full configured set:
npm run fixtures:prepare
npm run validateProbe hosted tool use and capture the machine record before paying for a matrix:
npm run connectivity
npm run environment -- --hash-modelCommit the finalized configuration, confirm a clean tracked worktree, and freeze it:
git add manifest.json cases graders examples scripts src test package.json package-lock.json
git commit -m "Freeze evaluation inputs"
npm run freeze
git add matrix-lock.json
git commit -m "Record matrix lock"Run an excluded smoke cell before the paid pilot:
npm run smoke -- --case ctx-uuid-validation --route gateway-deepseek-v4-flash
npm run pilotExecute or resume the matrix:
npm run matrixThe scheduler reads completed run records and skips cells already present. An excluded or incomplete cell must be rerun; it cannot silently enter the final matrix.
doctor → prepare → validate → connectivity → environment
→ freeze → smoke → pilot → matrix
→ reconcile → audit → report → analyze → methodology review
Each cases/<id>/case.json names two commits in a source repository:
baseCommit: the code shown to the agent;targetCommit: the historical change used to recover held-out tests and validate the grader;hiddenFiles: target-commit files copied into the grader workspace after the agent exits;benchmarkHiddenFiles: benchmark-authored oracles stored in this repository;hiddenCommands: required executable checks;validationCommands: broader required or diagnostic commands;expectedChangedPaths: the historical reference-patch footprint used for descriptive task-size slices, not an enforced allowlist;installCommand: dependency installation for the isolated fixture.
Prompts live beside the case configuration. They describe the requested behavior without revealing target implementation details or hidden assertions.
prepare verifies both commits, exports the base with git archive, initializes a new repository, removes access to remotes and future history, installs dependencies, and records the fixture commit. Working copies use APFS copy-on-write cloning on macOS and a recursive copy elsewhere.
Source repositories are read-only inputs. Fixture construction never checks out or modifies their working trees.
validate builds two grader workspaces:
- the untouched base, which must fail at least one required held-out command;
- the base plus the historical target patch, which must pass every required command.
A failed calibration blocks matrix execution. Diagnostic commands may record unrelated repository failures without deciding acceptance.
freeze records:
- the current harness commit;
- raw and environment-expanded manifest hashes;
- the exact schedule hash and cell count;
- SHA-256 values for engine, test, case, prompt, and grader files;
- the agent version and binary hash;
- configured file or Git artifacts, such as local weights and runtime source.
The matrix command checks the lock before every resumed run. Changes to an environment binding, scheduled route, grader, agent binary, or frozen artifact stop execution.
The child agent receives a fresh home directory and a minimal environment. Network tools, browser tools, subagents, Git history, remotes, and benchmark-owned graders are unavailable. The local relay rewrites the model name to the route’s pinned value and forwards the request.
Relay records contain request and response hashes, byte counts, token classes, timing, selected provider metadata, status, stop reason, cost, and timeout phase. Prompt text and completion text are omitted from relay logs.
The first patch is graded before any repair. When allowed, one repair pass receives only grader labels and bounded failure tails. The held-out files remain unavailable to the agent. Reports retain first-pass and final outcomes separately.
reconcile rebuilds a run’s relay totals from its append-only JSONL ledger. audit checks run completeness and consistency. report produces a compact CSV and summary. The included analysis adapter builds sanitized publication JSON for the local-versus-cloud experiment.
The review packet binds every displayed section with a SHA-256 value and redacts local roots. Configure the publication source in review/publication-files.json:
npm run methodology-review:packet -- --publication /path/to/publication-worktree
npm run methodology-review -- --publication /path/to/publication-worktreeReviewer requests and responses remain ignored under .state/reviews/.
manifest.json contains four groups:
- agent identity and environment controls;
- pilot, matrix, and repeated anchor schedules;
- process, network, output, and spend ceilings;
- local and hosted route definitions.
Machine-specific values use ${ENVIRONMENT_VARIABLE} placeholders. The loader refuses unresolved placeholders and never writes their resolved values back to the tracked manifest.
{
"agent": {
"adapter": "claude-code",
"binary": "${EVAL_AGENT_BINARY}",
"version": "2.1.221"
},
"environment": {
"requireAcPower": true
},
"routes": {
"local-model": {
"kind": "local",
"model": "fixed-model-id",
"upstreamBaseUrl": "http://127.0.0.1:8087",
"inputUsdPerMillion": 0,
"outputUsdPerMillion": 0,
"cacheReadUsdPerMillion": 0,
"cacheWriteUsdPerMillion": 0,
"runCapUsd": 0
},
"hosted-model": {
"kind": "gateway",
"model": "provider/model-id",
"upstreamBaseUrl": "https://ai-gateway.vercel.sh",
"inputUsdPerMillion": 1,
"outputUsdPerMillion": 5,
"cacheReadUsdPerMillion": 0.1,
"cacheWriteUsdPerMillion": 1.25,
"runCapUsd": 2
}
}
}Pricing in a manifest is experiment input. Verify current prices before freezing a new run. Exact billed metadata takes precedence over estimates when an endpoint supplies it.
Spend checks happen before a request or run is admitted. A final in-flight request can cross its request-admission threshold because its exact token usage and provider charge are known only after completion. Set reserves below the operator's absolute account budget.
fullPaidCapUsd accounts for every paid request made by the experiment workspace, including connectivity probes and excluded smoke or power-transition runs. Exclusion removes a cell from outcome analysis; it does not erase the provider charge. pilotPaidCapUsd applies only to cells launched by the pilot command.
See docs/CONFIGURATION.md for every field and docs/METHODOLOGY.md for scorer and inference guidance.
cases/ Historical task definitions and prompts
graders/ Benchmark-authored held-out oracles
examples/public-fixture Deterministic public base/reference calibration fixture
scripts/ Public-example generator and operator utilities
src/cli.mjs Command dispatcher
src/fixture.mjs Git export and isolated fixture creation
src/run.mjs Agent execution, scheduling, repair, restart logic
src/relay.mjs Protocol relay, telemetry, budgets, timeouts
src/grade.mjs Hidden-test overlay and patch-integrity scoring
src/freeze.mjs Matrix lock creation
src/lock.mjs Matrix lock verification
src/report.mjs Run reconciliation and normalized summaries
src/review.mjs Redacted evidence packets and reviewer execution
analysis/ Experiment-specific publication adapter
review/ Review models, rubric, and publication file map
test/ Offline relay, scheduler, grader, and privacy tests
.state/ Ignored fixture bases, runtime state, review packets
runs/ Ignored raw run records and agent traces
results/ Ignored generated summaries
Never commit runs/, .state/, .env, or generated result files without a deliberate disclosure review.
Raw agent logs can contain prompts, tool arguments, source fragments, and model output. Relay ledgers avoid body content by construction, yet provider metadata and error strings still require review. Public artifacts should come from the report, analysis, or audit-packet sanitizers.
The relay forwards requests to configured endpoints. A benchmark prompt that tells the agent to avoid network access does not control the upstream model provider. Review provider retention and routing policies separately.
- A route comparison evaluates complete deployed routes. It does not isolate one hardware or model variable.
- Historical public code may appear in model training even when target commits postdate a release.
- Single-run tasks reveal breadth; repeated anchors reveal variance for a smaller task subset.
- Sequential runs can inherit time-of-day, thermal, provider-load, and cache effects.
- Wilson intervals over cells do not model repeated-task dependence.
- A local API cost of zero excludes hardware amortization, electricity, engineering time, and opportunity cost.
- Provider fallback can improve reliability while changing latency, effective serving stack, or credential path.
- A deadline-terminated process can leave a passing captured patch. Process completion and patch correctness remain separate fields.
The engine has one completed 60-cell experiment, an independent methodology audit, a deterministic public calibration fixture, and a GLM 5.2 open-source release audit with final signoff. The public API and configuration schema may change before 1.0. Near-term work includes agent adapters, Linux power detection, JSON Schema validation, and a generic analysis layer. The repository remains private until the public release checklist is complete.
Read CONTRIBUTING.md before proposing changes. Security and disclosure reports belong in SECURITY.md.
MIT. See LICENSE.
