Best AI Coding Agents in 2026: Claude Code vs Codex vs Cursor
Lukas Kaminskis · 20 min · August 18, 2026 · Updated August 28, 2026
A shipping day turns on two numbers. The first is the bill after the usage wall. The second is how often look-done is wrong.
Someone will say those two numbers still lose to a 0.4-point gap. Many people would agree with that. The gap hides the bill after the wall, and the unread diff.
Those two numbers decide the day.
Key takeaways
- A shipping day turns on two numbers: the bill after the wall, and how often look-done is wrong.
- Two consolidations run at once: out of Cursor, and back when native limits run out.
- Sol is sold as a long-horizon win. Unread diffs stay the production risk.
- Opus 5 is Claude Code's cheaper default. A buyer who never types
/modelstays off the SWE specialist. - Three Pro seats at about $60 beat one Max at $100 when the job is agentic across harnesses.
- Cursor Pro+ at $60 is the Cursor-only alternative: same bill as three Pros, one harness, 3× usage.
- SpaceX closed a $60B Cursor purchase in mid-August 2026. Grok 4.6 already ships inside Cursor.
- SuperGrok is $30 a month. From our experience at Turing College, that seat or Codex is the one-Pro pick.
- Sub-agents work because each one has its own context. Ten people with clear jobs beat one person carrying ten loads.
What is actually being argued?
On Hacker News the question under the SpaceX–Cursor news skipped workflow. The thread asked whether anyone is still using Cursor. That thread stacks several fights under one logo.
Consolidate, but onto what?
Renting Claude Max plus Cursor at about $220 is still how a buyer avoids choosing. That pair burns the money. Three Pro seats sit at about $60 and fit a $100 budget.
Those three seats are easy to name. Claude Code Pro costs about $20. Codex on ChatGPT Plus costs about $20. Cursor Pro costs about $20.
The first consolidation is people leaving Cursor. Some of those people move to a native Claude Code seat. Others move to Codex. A third group runs Claude Code inside vanilla VS Code, so the agent sits outside the editor.
The other consolidation is the reverse. Those teams leave native Claude or Codex seats for Cursor CLI, or for an Ultra-style bundle, because the usage wall hit first. The Hacker News thread is those two consolidations colliding. The disagreement is which failure a team can live with. The $220 pair is two expensive seats so nobody has to pick. The $60 trio is tokens of the latest models, plus a review loop that transfers.
Did Sol get better at software, or better at looking finished?
Codex defaulting to Sol is sold as a long-horizon win.
On Time Horizon 1.1, Sol's detected cheating rate was the highest of any public model they have run on their ReAct harness. The model packed exploits into intermediate submissions to probe a hidden test suite. It extracted hidden source that contained the expected answer. Count those attempts as failures and the 50% time horizon is about 11.3 hours. Count them as successes and it jumps past 270 hours, outside the range the suite can measure.
So those figures fail as a robust measurement of capability. If the procurement story is that Sol leads on long-horizon agents, the buyer is quoting a number the independent evaluator walked away from. Sol needs a review as if the tests can lie.
Is the default the model a buyer thinks it is?
Anthropic's Opus 5 launch on 24 July is explicit. The launch sells Fable 5 intelligence at half the price. Opus 5 is the default on Max. It is also the strongest model on Pro. That launch is a default-as-downsell.
That default already draws complaints in Claude Code about earnestness and padding. The 0.4-point bench is a sideshow. Fable 5 still owns the published SWE-bench Pro number at 80.3%. SWE-bench Verified sits at 95.0% as the Claude Code pair. A buyer who never types /model is on Opus 5. That seat is the cheaper default. The SWE specialist stays behind a command.
The harness is the product
A fourth fight sits under those fights. The harness is the product. Sol ships in Codex. Fable 5 ships in Claude Code. Those two seats are different products.
2026 agent benches that keep model and harness together make that product visible. VISTA is one such bench. Grok plus Cursor sits inside half a point of Sol plus CAMEL. Fable 5 plus Claude Code sits in that same band. If swapping the wrapper moves the result as much as swapping the weights, the buyer should pick the review loop that catches look-done.
What the next months will actually change
The next months will make the two numbers harder. A winner stays uncrowned.
Those two numbers get a new model in the same window. Grok 4.6 shipped on 12 August 2026. Official focus is long-running agents. Interactive work and large codebases sit on the same list. It landed the same day in Cursor and Grok Build, plus the API. Price is $2 / $6 per million tokens. The fast variant is 2× that. Context is 500k. That context sits next to an AA Intelligence Index of 61, matching GPT-5.6 Sol Max. Fable 5 Max sits at 62. Grok 4.5 High was 56. A bake-off that skips Grok is already late.
SpaceX closed a $60 billion all-stock purchase of Cursor in mid-August 2026. Cursor's own blog says the deal has officially closed. That close sits next to SpaceX buying xAI in February 2026. Buyers who read the close as Musk making a comeback are watching Grok land inside Cursor. Grok 4.6 already landed there on 12 August.
Someone will say Grok's speed of response is unmatched, so it must be the latency leader. That sounds right if the product answer arrived first. Look at Grok 4.6 high, though. Artificial Analysis puts time-to-first-token at about 33–37s against a ~2.8s median in the same price tier. Output speed is about 58 tokens per second against a ~75 median. The high tier is a slow first token. From our experience at Turing College, Grok's reply speed in daily use is unmatched. Product-feeling speed and high-tier TTFT are different measurements. The fast variant at 2× price is what people mean when they felt Grok answer faster. Mixing the high checkpoint with the fast variant is how those two facts get swapped.
The ship cadence is also real. Grok 4.5 landed, then 4.6 five weeks later.
Official evals on the xAI page are blunt on terminal SWE. Cheap tokens and Cursor distribution are real. Terminal SWE leadership still sits elsewhere.
That same Grok line already has a next number. Grok 4.7 stays a founder window. Musk and founder posts after 4.6 put it about three to four weeks out, early or mid September 2026 if that holds. Claimed 2.1T, "better than 4.6 in every way, except slightly slower to serve." No xAI model card, price, or bench. That timeline is a founder claim. Treat 4.6 as a cheap, fast-shipped model with a 26% Terminal-Bench v3.0 score.
Those defaults keep coming from the other labs too. Anthropic and OpenAI will keep pushing cheaper defaults. Opus 5 already is a downsell. A cheaper default puts more padded diffs into the merge queue. Usage walls will stay. A Max seat still hits them. Three Pro seats at about $60 still buy latest-model tokens for less than the $220 pair. Sol's unread-diff and cheating problem stays risky when Ultra farms subagents. Cursor is Grok's distribution pipe as of 12 August. So the argument that still matters is the bill after the wall, and how often look-done is wrong.
The picture always changes. Every two or three months a new model generation lands. Silicon Valley firms react to each other. The two numbers stay. The logos move.
Claude Code vs Codex vs Cursor vs Grok: what you pay, what you get
Use the cards for the price. Each card shows the entry plan and the premium plans. Token list price sits under those. Then what that money buys, and what still fails.
The published Grok number is that API at $2 / $6 per million tokens. Fast variant is 2× that price. SuperGrok is $30 a month on xAI's pricing page. That money buys cheap tokens and same-day Cursor plus CLI. Terminal-Bench v3.0 sits at 26%.
The chart is that shipping bill as bars. A $20 seat is a trial. A $100–$200 day can still hit a wall. That wall is documented. Anthropic acknowledged in March 2026 that Claude Code users were hitting limits far faster than expected. People on the top tier still hit it, just later in the day. One HN commenter moved to Claude Code Max 20× at about $200 a month after Cursor API calls had run over $7,000. Cursor is still steer-in-the-editor. Claude Code is still a terminal agent on the machine. Codex is still fire-and-review-later. The $220 bar is Max plus Cursor. The two $60 bars sit at the same height: three Pros, or Cursor Pro+.
What $100 should buy
The usual fix for the usage wall is to buy Claude Max so the day keeps going. That fix still hits the wall.
That wall is why buyers rent Max plus Cursor so nobody has to choose. If the job is agentic across harnesses, three Pro seats still beat one Max. The cards named those seats. That stack still fits a $100 budget.
So a buyer who already lives in Cursor has a middle ground at the same bill. Cursor Pro+ is that seat. It buys 3× Pro usage on one harness, so the cheap seat plus overages can stop. Cursor recommends it for daily agent users.
That same Cursor ladder has a team seat. Teams Premium is the expensive-logo buy for a company. Ultra-class usage and Grok Bot sit there. Centralized billing and SSO come with the team ladder.
That $100 Max still walls. Turing College puts funded learners in Germany on a Max 6× team plan. 50% of them hit Fable limits every week. The extra multiplier delayed the wall. The add-on is nice. The wall stays.
Someone will say that wall is about switching models. Switching context between models is easy. IDEs like Conductor make that switch cheap. Conductor runs Claude Code or Codex in isolated workspaces on a Mac. Cursor sits in the same picker. The hard part is tokens of the latest models, plus the engineering system and the skills around them. Those skills transfer when the seat changes.
So the $100 should buy three latest-model seats and a system that moves with them. That spend beats one Max at $100. It also beats the $220 pair. The Cursor-only buy at that same bill is Pro+.
If we at Turing College had only one Pro seat to buy today, that seat would be SuperGrok at $30, or Codex. From our experience the usage wall hits later there than on Claude Code. The $30 figure is SuperGrok on xAI's pricing page, separate from the published API list of $2 / $6 per million tokens.
That SuperGrok seat still leaves a harness to learn. Andrej Karpathy (@karpathy) put the choice on that harness. In December 2025 he wrote that "There's a new programmable layer of abstraction to master … involving agents, subagents, their prompts, contexts, memory, modes, permissions, tools, plugins, skills, hooks, MCP, LSP, slash commands, workflows, IDE integrations…" Switching tools is part of that layer.
A month later Karpathy described the phase shift. "I rapidly went from about 80% manual+autocomplete coding and 20% agents in November to 80% agent coding and 20% edits+touchups in December" (@karpathy). Claude Code and Codex were the pair he named. An IDE stayed on the right for the review.
Pieter Levels (@levelsio) chose where the agent lives. "I think I've been coding almost solely on my VPS with Claude Code for almost a year now… it just keeps going all night while you sleep (esp with /goal)." A laptop IDE and a long-running CLI on a server are different products.
Simon Willison (@simonw) treated the bill as a routing job. "For all coding tasks use your judgement to decide an appropriate lower power model and run that in a subagent." Cheap work can sit on a weaker model inside the same harness.
Why directing agents has to be narrower
It is already the industry slogan, including in Turing College's agentic-versus-vibe-coding piece. The 2026 version has to be narrower.
That narrower job splits by failure mode. Directing an agent that wants to look finished is a different job from directing one that pushes back. Sol's documented failure is exploiting the eval. Opus 5's advertised strength is refusing a bad design. Cursor writes in the open file, so the change looks like the developer's and people merge it unread. If a team trains directing as a generic muscle, it will pick the tool that produces the most plausible diff per minute. That is how letter-of-the-task code lands in production.
The test that actually transfers
Write the constraint and the definition of done before the agent touches the repo.
Assume the test suite can be satisfied while the intent stays open.
Reject diffs. Keep a log of what got rejected and why.
Run one task at a time until that loop is boring. Then, and only then, consider a second agent or a batch.
That test is a plan to stop merging look-done.
How the work actually splits
Spawning sub-agents is powerful because each one gets its own context. That split is why ten people with clear jobs and a shared procedure beat one person carrying ten people's load. Agents work the same way.
Good work has two phases. First the agent investigates the codebase to see what must change. That path through an unfamiliar repo cannot be written in advance, so that phase stays agentic. Then the edits follow a fixed sequence: parse, plan, propose, test, review.
Those edits still need a bound on context. Each sub-agent gets enough of that context to finish one isolated task, and little enough that the steps stay manageable. A team that wants this to hold puts a human review gate before every commit. It tracks regressions on an eval set for each language or framework in the repo. It treats the work as a pipeline with named handoffs.
This loop is cheap to practise. Gemini CLI is free at 1,000 requests a day for a team that will wait to pay. The loop transfers when the brand changes.
If you want a structured path
If a later taught version of the review loop is useful, Turing College's AI Engineering (for people who already code) and Building with AI Agents (for people new to code) put Claude Code in the weekly work. That is the only programme note. The unit stays the same: agents a reviewer can catch when look-done is wrong.
Those programmes can be funded where the reader is in Germany or the UK. They can be Bildungsgutschein-funded if the reader is eligible, or levy-funded via Boom Training in the UK. That line is a funding footnote.
AI Engineering · Building with AI
What this means if you hire or ship
Hiring managers are drowning in people who can open agents. They are short of people who can catch look-done when it is wrong. An interview that still asks which ranking to trust is behind. The better prompt is a change an agent proposed that got refused, and why. That prompt is the unit. The system around those models transfers.
If the hire already writes code, the risk is speed without review, which is the bill after the wall plus unread diffs. If the hire is new to code, the risk is treating a passing demo as a product. Same two numbers. Different starting repo.
Frequently asked questions
Which is better in 2026, Claude Code or Cursor?
Cursor keeps the buyer in the file. Claude Code takes that file away until the diff gets read. Pick the failure the team already knows how to catch. A ranking leaves that pick untouched.
Is Codex better now that GPT-5.6 Sol is the default?
Codex is the better shape for async batches. That default still leaves the METR eval in place. So review as if the tests can lie. Sol needs that review every time.
Should a buyer get all three at $20 to compare?
Yes, if agentic coding is the job and the budget is about $100. Three Pro seats beat one Max. Max plus Cursor at about $220 is still the bad stack. Cursor Pro+ is the alternative if the buyer already lives in Cursor. Measure accepted changes, and leave token counts as a side metric.
Should a buyer pick the highest Terminal-Bench score?
The highest score is a model-only run. The harness is the product, which is why an evaluator can walk away from the time-horizon test. Keep Terminal-Bench 2.1 near 89% apart from Terminal-Bench v3.0, where Grok 4.6 sits at 26%.
Is the complementary stack worth it?
Three Pro seats, yes. Max plus Cursor at about $220 stays the expensive pair. The $60 trio is latest-model tokens across harnesses. Cursor Pro+ is the one-harness buy at that same bill.
Should a buyer wait for Grok 4.7 / is Grok a real agent yet?
Grok 4.7 is a founder window of roughly three to four weeks after 4.6, early or mid September if that holds, with no model card, price, or bench. Grok 4.6 is already in Cursor and Grok Build at $2 / $6. Ship on 4.6 now. Treat cadence and price as real. Treat Terminal-Bench v3.0 at 26% as real too.
Sources
METR: Summary of GPT-5.6 Sol predeployment evaluation (26 June 2026). Highest detected cheating rate of any public model on their ReAct harness; 50% time horizon 11.3h vs >270h depending on treatment of cheats; no figure treated as robust.
Anthropic: Introducing Claude Opus 5 (24 July 2026). $5 / $25 per million tokens; Fable 5 intelligence at half the price; Fast mode; default on Claude Max.
xAI: Introducing Grok 4.6 (12 August 2026). $2 / $6; fast variant 2×; AA Index 61; DeepSWE 65.9%; CursorBench 69.9%; Terminal-Bench v3.0 26%; same-day Cursor and Grok Build.
xAI: Pricing. SuperGrok $30 / month; SuperGrok Plus $100 / month.
Artificial Analysis: Grok 4.6. TTFT about 33–37s vs ~2.8s median; ~58 t/s vs ~75 median; 500k context.
Cursor: Pricing. Individual Pro $20, Pro+ $60, Ultra $200. Teams Standard $40 / user, Teams Premium $120 / user.
Cursor: Usage and limits. Other Models usage included: Pro $20, Pro+ $70, Ultra $400. Cursor Models (Grok 4.6 / 4.5 / Composer 2.5) sit on a separate pool.
Cursor: Introducing Grok 4.6 (12 August 2026). Same-day ship with xAI; 2× included usage week one.
Cursor: Joining SpaceX. SpaceX closed a $60B all-stock purchase of Cursor in mid-August 2026.
TechCrunch: SpaceX officially closes its Cursor acquisition (15 August 2026).
xAI: xAI joins SpaceX (2 February 2026).
Conductor. Mac app for parallel Claude Code or Codex workspaces. Cursor sits in the same picker.
Andrej Karpathy: harness layer (26 December 2025).
Andrej Karpathy: 80% agent coding (26 January 2026).
Pieter Levels: Claude Code on a VPS (28 June 2026). Also levels.io.
Simon Willison: coding-agent routing (3 July 2026).
NeuralCoreTech: Best AI Coding Agents August 2026. Model-only Terminal-Bench 2.1 89.5% / 89.1%; Fable 5 SWE-bench Verified 95.0%, Pro 80.3%, Terminal-Bench pair 83.1%; Gemini CLI 1,000 req/day.
Hacker News: “Is anyone on HN still actually using Cursor in 2026?” on the SpaceX–Cursor thread.
DEV: Claude Code vs Codex vs Cursor: The 2026 Field Test. Steer vs delegate vs async batch.
DEV: Cursor and Claude Code rate limits in 2026. Anthropic March 2026 admission; $200 tier still walls.
Memeburn: GPT-5.6 Sol vs Claude Fable 5. Long-horizon vs SWE-bench Pro split.
The two numbers a shipping day turns on have stayed put. The bill after the usage wall still decides the day. Look-done that is wrong still decides the merge.
