Skip to main content
Applied AI for Management

Claude Code, Codex or Antigravity? Why Scrum Masters Should Not Choose on Benchmark Scores

Benchmark tables score the model, not the tool, and the best-known coding benchmark has been dropped by OpenAI. Four questions make a better recommendation: which plan, where your instructions live, whether you can read the code, and whether your organisation allows it.

Three glowing cubes in blue, amber and purple feed into a blueprint of structural diagrams on a desk, with faded bar charts dissolving in the background.

If you have been asked to evaluate Claude Code, OpenAI's Codex and Google's Antigravity, you will be offered a benchmark table within minutes of starting. It feels like the defensible basis for a recommendation. It is the weakest one available. Benchmark scores measure the model behind a tool rather than the tool, they move every few months, and the best-known coding benchmark has been abandoned by one of the companies that used to report it. This post gives you four questions that make a better recommendation, with the answers for all three tools, checked at source on 2 October 2026.

The clearest sign that benchmarks are the wrong axis comes from one of the vendors. OpenAI published a post with a title that says it plainly:

"Why we no longer evaluate SWE-bench Verified" OpenAI

SWE-bench Verified was the coding benchmark most comparison tables quoted. OpenAI stopped evaluating on it after concluding it had become contaminated, with frontier models able to reproduce some of its answers from training data, as reported at the time. Any comparison still built on it is comparing something other than what you think.

Why benchmarks miss the point twice over

First, a benchmark scores a model. These tools are not models. They are the program around the model: how it reads your files, what it asks before acting, where your instructions live, and which subscription it draws from. Google's own Antigravity pricing page lists Claude Sonnet and Opus among the models available inside Antigravity. The same model can sit inside more than one of these tools, which makes "which tool scores higher" a confused question.

Second, the benchmarks test writing code. A Scrum Master or project manager is mostly asking for summaries, reports, structured documents and checks against a board. A coding score says little about that work.

Two of those answers matter more than they look. The plan question decides what your usage competes with: on Claude, a heavy Claude Code week also spends your chat allowance. And the instructions question decides how locked in you are. Since version 2.1.277, Claude Code can read an AGENTS.md file when there is no CLAUDE.md, and AGENTS.md is the file Codex uses. Instructions written once can serve both.

"My manager wants a number. A benchmark is defensible."

Nobody gets questioned for recommending the tool that scored highest. A list of structural trade-offs sounds like an opinion.

A benchmark is defensible until someone asks which model the score was for, whether it was the model your organisation can actually use, and whether the benchmark is still trusted. Each of those questions is now awkward. The four questions above are answered from the vendors' own pages, which is a stronger footing than a third-party table.

They also answer what the person asking actually needs to know: what it will cost, what it will compete with, and how hard it would be to leave.

"Surely the fastest one is the best one."

If one of them answers four times faster, that is hours back every week.

Speed figures in comparison articles measure how fast a model produces text, under one load, on one day. They depend on the model chosen inside the tool, not just the tool, and they move as vendors change capacity. We have not quoted any here because we could not confirm one at source.

More to the point, for reporting work the bottleneck is rarely generation speed. It is checking the output. A summary that arrives in five seconds and needs two corrections is slower than one that takes thirty seconds and needs none.

"Whatever I build in one, I'm stuck with."

If I spend a month building skills and instructions in one tool, I cannot move without starting again.

Less than you might fear. Instructions written in plain text in an AGENTS.md file can be read by Codex and, from version 2.1.277, by Claude Code. Skills in Claude Code are plain-text files of instructions, so even where a format differs, the substance of what you wrote, the rules of your sprint report and the structure of your stakeholder update, is readable and movable.

That is the practical meaning of portability. Your judgment, written down in plain language, is the asset. The tool that reads it is replaceable.

The part worth keeping

The benchmark table answers a question nobody evaluating these tools for project work should be asking. Ask what each one costs and shares, where your instructions live, whether you can inspect it, and whether your organisation allows it. Those answers come from the vendors themselves, they decide your real experience, and they make a recommendation you can defend in a meeting.

Sources: OpenAI, Why we no longer evaluate SWE-bench Verified | Google, Antigravity pricing | OpenAI, Codex repository (plans and licence) | OpenAI, Codex AGENTS.md | Anthropic, Claude Code memory and AGENTS.md | Anthropic, Claude Code with Pro or Max. Checked 2 October 2026.

AI Systems with Claude, for Scrum Masters and Project Managers

Most AI pilots in project management do not fail on the model. They fail because somebody pointed the thing at a decision instead of at a task. This course is built around that distinction, and around the work that follows once you get it right.

Seven sections, taught against one running project from the first lecture to the last. You watch a system get built, then you build the same one against your own sprint. Nothing here is a demo that works only on the example.

You finish with five working systems and you keep them. The sprint report machine turns your board export and standup notes into the report you currently write by hand. The retro intelligence system tells you what the team keeps saying, not just what it said this time. The stakeholder comms engine drafts the update in the register the audience expects. The backlog health monitor flags the quiet decay nobody has time to check for. The risk and dependency tracker follows the chains that actually bite.

Seventy-six working files ship with it, across four sections, so every system is built against real sprint data rather than material invented for a slide. Four sprints of retrospectives, board exports, standup notes, backlog and risk data.

Explore the Course


Build the Sprint Reporting System, Not Another Prompt

AI Systems with Claude is our own course for Scrum Masters and Agile Project Managers, taught against one running project across seven sections. You build five working systems and keep them: the sprint report machine, the retro intelligence system, the stakeholder comms engine, the backlog health monitor, and the risk and dependency tracker. Direct from HK School of Management, not on Udemy.

See What Is Included