Model Benchmarks
Performance of coding agents creating and changing eve projects, measured by deterministic checks against the files, commands, and simulated external interactions each run produces1.
Last published: October 8, 2026eve revision: 279e2b6a
Current results
| Claude Code | 218.6s | $1.03 | 100% | 100% | |
| Claude Code | 152.8s | $0.33 | 100% | 100% | |
| Claude Code | 122.4s | $0.29 | 100% | 100% | |
| OpenCode | 319.6s | $0.09 | 90% | 100% | |
| Codex | 125.4s | $0.16 | 100% | 100% | |
| Codex | 159.9s | $0.86 | 100% | 100% | |
| Codex | 157.4s | $0.22 | 100% | 100% | |
| OpenCode | 163.1s | $0.14 | 86% | 100% | |
| OpenCode | 176.3s | $0.08 | 90% | 100% | |
| OpenCode | 160.5s | $0.17 | 100% | 100% | |
| OpenCode | 141.3s | $0.03 | 90% | 100% | |
| Codex | 151.9s | $0.01 | 95% | 95% | |
| OpenCode | 176.3s | $0.22 | 86% | 90% | |
| OpenCode | 302.3s | $0.43 | 100% | 90% |
Retired Models
| Claude Code | 110.3s | $4.37 | 100% | 100% | |
| Claude Code | 139.1s | $0.86 | 100% | 100% | |
| Codex | 122.7s | $0.23 | 100% | 100% | |
| OpenCode | 77.4s | $0.14 | 86% | 100% | |
| OpenCode | 184.9s | $0.11 | 62% | 90% | |
| OpenCode | 233.1s | $0.07 | 19% | 38% |
1 Each model completes the same deterministic eve authoring tasks three times with and without the guidance files generated by eve init; success rates include every scheduled run, while duration, tokens, and tool calls average valid runs. Cost estimates apply provider public list prices to recorded token usage.