Agent Evaluations
Performance of coding agents creating and changing eve projects, measured by deterministic checks against the files, commands, and simulated external interactions each run produces.
| OpenCode | 102.8s | $0.20 | 100% | 100% | |
| OpenCode | 100.4s | $0.08 | 86% | 100% | |
| OpenCode | 94.6s | $0.15 | 48% | 100% | |
| OpenCode | 57.1s | $0.02 | 86% | 100% | |
| OpenCode | 58.8s | $0.06 | 95% | 100% | |
| OpenCode | 101.4s | $0.12 | 100% | 100% | |
| OpenCode | 48.1s | $0.05 | 90% | 95% | |
| OpenCode | 110.3s | $0.09 | 71% | 81% | |
| OpenCode | 233.1s | $0.07 | 19% | 38% |
* eve guidance includes the AGENTS.md and agent-specific aliases generated by eve init. The column shows the success rate when agents have access to this guidance.
N/A marks a case that a model has not run. It counts against that model's success rate so every model is scored against the same case set. When a case changes, eve retains a model's most recent result until that model reruns it.
Avg List Cost is the mean estimated cost per evaluation from each run's token usage at the provider's public list price. It is a relative comparison, not a Gateway bill: routing, cache discounts, and provider pricing can differ.