If you have been asked to evaluate Claude Code, OpenAI's Codex and Google's Antigravity, you will be offered a benchmark table within minutes of starting. It feels like the defensible basis for a recommendation. It is the weakest one available. Benchmark scores measure the model behind a tool rather than the tool, they move every few months, and the best-known coding benchmark has been abandoned by one of the companies that used to report it. This post gives you four questions that make a better recommendation, with the answers for all three tools, checked at source on 2 October 2026.
The clearest sign that benchmarks are the wrong axis comes from one of the vendors. OpenAI published a post with a title that says it plainly:
"Why we no longer evaluate SWE-bench Verified" OpenAI
SWE-bench Verified was the coding benchmark most comparison tables quoted. OpenAI stopped evaluating on it after concluding it had become contaminated, with frontier models able to reproduce some of its answers from training data, as reported at the time. Any comparison still built on it is comparing something other than what you think.
Why benchmarks miss the point twice over
First, a benchmark scores a model. These tools are not models. They are the program around the model: how it reads your files, what it asks before acting, where your instructions live, and which subscription it draws from. Google's own Antigravity pricing page lists Claude Sonnet and Opus among the models available inside Antigravity. The same model can sit inside more than one of these tools, which makes "which tool scores higher" a confused question.
Second, the benchmarks test writing code. A Scrum Master or project manager is mostly asking for summaries, reports, structured documents and checks against a board. A coding score says little about that work.
Two of those answers matter more than they look. The plan question decides what your usage competes with: on Claude, a heavy Claude Code week also spends your chat allowance. And the instructions question decides how locked in you are. Since version 2.1.277, Claude Code can read an AGENTS.md file when there is no CLAUDE.md, and AGENTS.md is the file Codex uses. Instructions written once can serve both.
"My manager wants a number. A benchmark is defensible."
Nobody gets questioned for recommending the tool that scored highest. A list of structural trade-offs sounds like an opinion.
A benchmark is defensible until someone asks which model the score was for, whether it was the model your organisation can actually use, and whether the benchmark is still trusted. Each of those questions is now awkward. The four questions above are answered from the vendors' own pages, which is a stronger footing than a third-party table.
They also answer what the person asking actually needs to know: what it will cost, what it will compete with, and how hard it would be to leave.
"Surely the fastest one is the best one."
If one of them answers four times faster, that is hours back every week.
Speed figures in comparison articles measure how fast a model produces text, under one load, on one day. They depend on the model chosen inside the tool, not just the tool, and they move as vendors change capacity. We have not quoted any here because we could not confirm one at source.
More to the point, for reporting work the bottleneck is rarely generation speed. It is checking the output. A summary that arrives in five seconds and needs two corrections is slower than one that takes thirty seconds and needs none.
"Whatever I build in one, I'm stuck with."
If I spend a month building skills and instructions in one tool, I cannot move without starting again.
Less than you might fear. Instructions written in plain text in an AGENTS.md file can be read by Codex and, from version 2.1.277, by Claude Code. Skills in Claude Code are plain-text files of instructions, so even where a format differs, the substance of what you wrote, the rules of your sprint report and the structure of your stakeholder update, is readable and movable.
That is the practical meaning of portability. Your judgment, written down in plain language, is the asset. The tool that reads it is replaceable.
The part worth keeping
The benchmark table answers a question nobody evaluating these tools for project work should be asking. Ask what each one costs and shares, where your instructions live, whether you can inspect it, and whether your organisation allows it. Those answers come from the vendors themselves, they decide your real experience, and they make a recommendation you can defend in a meeting.
Sources: OpenAI, Why we no longer evaluate SWE-bench Verified | Google, Antigravity pricing | OpenAI, Codex repository (plans and licence) | OpenAI, Codex AGENTS.md | Anthropic, Claude Code memory and AGENTS.md | Anthropic, Claude Code with Pro or Max. Checked 2 October 2026.


