Under construction. This evaluation suite is actively under development. Its results, methodology, and presentation are subject to change.

Agent Evaluations

Performance of coding agents creating and changing eve projects, measured by deterministic checks against the files, commands, and simulated external interactions each run produces.

Last published: August 19, 2026eve revision: 54a028ba
OpenCode102.8s$0.20100%100%
OpenCode100.4s$0.0886%100%
OpenCode94.6s$0.1548%100%
OpenCode57.1s$0.0286%100%
OpenCode58.8s$0.0695%100%
OpenCode101.4s$0.12100%100%
OpenCode48.1s$0.0590%95%
OpenCode110.3s$0.0971%81%
OpenCode233.1s$0.0719%38%

* eve guidance includes the AGENTS.md and agent-specific aliases generated by eve init. The column shows the success rate when agents have access to this guidance.

N/A marks a case that a model has not run. It counts against that model's success rate so every model is scored against the same case set. When a case changes, eve retains a model's most recent result until that model reruns it.

Avg List Cost is the mean estimated cost per evaluation from each run's token usage at the provider's public list price. It is a relative comparison, not a Gateway bill: routing, cache discounts, and provider pricing can differ.