Model Benchmarks
Performance of coding agents creating and changing eve projects, measured by deterministic checks against the files, commands, and simulated external interactions each run produces1.
Last published: September 2, 2026eve revision: 78fa9046
Current results
| Claude Code | 154.4s | $1.01 | 100% | 100% | |
| Claude Code | 139.1s | $0.86 | 100% | 100% | |
| Claude Code | 125.2s | $0.35 | 100% | 100% | |
| OpenCode | 195.8s | $0.08 | 100% | 100% | |
| Codex | 122.7s | $0.23 | 100% | 100% | |
| Codex | 110.4s | $0.15 | 100% | 100% | |
| OpenCode | 77.4s | $0.14 | 86% | 100% | |
| OpenCode | 117.7s | $0.11 | 86% | 100% | |
| OpenCode | 95.5s | $0.21 | 57% | 71% |
Retired Models
| Claude Code | 110.3s | $4.37 | 100% | 100% | |
| OpenCode | 124.5s | $0.06 | 100% | 81% | |
| OpenCode | 233.1s | $0.07 | 19% | 38% |
1 Each model completes the same deterministic eve authoring tasks three times with and without the guidance files generated by eve init; success rates include every scheduled run, while duration, tokens, and tool calls average valid runs. Cost estimates apply provider public list prices to recorded token usage.