Live Comparison / Multi-Model
Compare which models repaired verifier failures under the same budget.
Every model here saw the same six tools, the same verifier loop, and the same four-call budget. The leaderboard shows who converged, how quickly they converged, and where they ran out of room.
Ranking
Model leaderboard
Every entry below comes from the same six-tool suite and the same four-iteration budget per tool.
Per-tool outcomes
Where each model converged or exhausted budget
| Model | Converged | Avg iterations | Per-tool result |
|---|