Skip to content
🚀 AI App Builders

Build full-stack apps from natural language

LovableReplitBolt.newBase44V0 by VercelBrowse All →
💻 AI Coding Tools

Code faster with AI pair programmers

GitHub CopilotCursorWindsurfTabnineCodexClaude CodeBrowse All →
🔍 AI Research Tools

Find and synthesize research faster

PerplexityChatGPTConsensusElicitBrowse All →
🎙️ AI Meeting Assistants

Record, transcribe, and summarize meetings

Fireflies.aiFathomOtter.aiGranolaBrowse All →
Home/AI Coding Tools/Codex vs Claude Code
CodexCodexIncluded with eligible ChatGPT plans; flexible credits and enterprise options are available
VS
Claude CodeClaude CodeIncluded with eligible Claude plans; Console, API, Team, and Enterprise access are also available

Codex vs Claude Code: 108-Run Coding Benchmark

We tested both coding agents on the same 18 unpublished TypeScript, Python, and Go tasks, with three clean repetitions per tool and no human steering.

Research curated by Etienne Tawong · Verified July 2026 · Our methodology →

Verified ResultLast verified: July 25, 2026
Provisional Winner
Codex
87.0
out of 100 (normalized)
87 / 100 raw points
vs
Claude Code
61.1
out of 100 (normalized)
61.1 / 100 raw points

Codex won the pre-registered primary outcome at 87.0 percent autonomous success versus Claude Code at 61.1 percent. The task-clustered 95 percent confidence interval for the difference ran from 11.1 to 42.6 percentage points in Codex's favor.

Directly verifiedPlatform reportedNot independently verified

Directly verified labels mark findings backed by the frozen harness, logs, patches, and evaluator results.

Primary Outcome

CategoryWeightCodexClaude Code
Autonomous task success1008761.1
Raw total1008761.1
Normalized (/100)87.061.1

The page score is the preregistered primary outcome: autonomous task success across 54 canonical runs per tool.

How We Tested

  1. The controlled matrix used 18 unpublished tasks: six TypeScript, six Python, and six Go. Each tool ran every task three times from a clean repository snapshot, producing 108 canonical runs.
  2. Both tools received identical prompts, public tests, hidden evaluators, caches, machine resources, and wall-clock limits. The schedule was randomized once from the frozen seed 20260719.
  3. Vendor-specific extensions were disabled. There was no web search, browser access, MCP, plugins, project-specific instructions, subagents, cross-session history, or human steering after prompt submission.
  4. Codex CLI 0.144.6 used gpt-5.6-sol. Claude Code 2.1.215 used claude-opus-4-8 at maximum effort. Shell network access was blocked except for each product's control plane.
  5. Autonomous success required a clean completion, a buildable patch, the full public regression suite, and the complete adjudicated hidden suite. This is stricter than the frozen 90 percent hidden-criteria threshold.
  6. Statistics use paired task evidence. Confidence intervals resample whole tasks so the three repetitions of one task are not treated as independent tasks.

What Each Tool Produced

Screenshots captured during the run. Click any image to view it full size.

Directly verified

Codex completed a TypeScript preferences task with the public and hidden evaluators passing.

Directly verified

Claude Code completed the paired TypeScript preferences task and returned a passing implementation.

Directly verified

Codex completed a Python cursor task under the frozen harness without human intervention.

Directly verified

Claude Code completed the paired Python cursor task under the same harness and limits.

Directly verified

Codex completed a Go configuration task with the public and hidden evaluators passing.

Directly verified

Claude Code completed the paired Go configuration task with the same evaluator setup.

Direct Observations

  • Codex completed 47 of 54 runs autonomously, or 87.0 percent. Claude Code completed 33 of 54, or 61.1 percent.
  • The success-rate difference was 25.9 percentage points. Its task-clustered 95 percent bootstrap interval was 11.1 to 42.6 points in Codex's favor.
  • Codex won 11 tasks, Claude Code won 1, and 6 were ties. The task-level two-sided exact sign-test p-value was 0.006348.
  • Codex passed 148 of 156 adjudicated hidden test cases, or 94.9 percent. Claude Code passed 135 of 156, or 86.5 percent.
  • Codex's median completion time was 83.5 seconds. Claude Code's was 357.6 seconds, about 4.3 times longer.
  • Codex had no provider or CLI errors in 54 canonical runs. Claude Code had 10. Quota exhaustion was retained as an outcome under the frozen protocol.
  • Without any adjudications, Codex still led on autonomous success at 72.2 percent versus 53.7 percent.

Functional Results

Each finding below is backed by the frozen run corpus and objective evaluator output.

Primary outcomeDirectly verified
  • Codex: 47 of 54 autonomous successes, 87.0 percent
  • Claude Code: 33 of 54 autonomous successes, 61.1 percent
  • Difference: 25.9 percentage points for Codex
  • Task-clustered 95 percent interval: 11.1 to 42.6 points
Correctness and regression evidenceDirectly verified
  • Codex: 148 of 156 adjudicated hidden test cases passed
  • Claude Code: 135 of 156 adjudicated hidden test cases passed
  • Both tools passed all 54 public regression suites
  • Raw hidden results also favored Codex, 89.7 percent to 80.1 percent
Reliability and speedDirectly verified
  • Codex median completion time: 83.5 seconds
  • Claude Code median completion time: 357.6 seconds
  • Codex provider or CLI errors: 0 of 54
  • Claude Code provider or CLI errors: 10 of 54

Limitations Of This Benchmark

  • This is the controlled CLI track. Native product features such as cloud tasks, editor integrations, hooks, skills, subagents, and MCP were deliberately disabled and were not scored.
  • The result is a July 2026 snapshot of specific tool versions, model identifiers, subscription accounts, and one Apple silicon machine.
  • Provider limits affected Claude Code during the run window. They were retained because the protocol preregistered quota exhaustion and tool errors as outcomes.
  • Five adjudication rules addressed unstated evaluator assumptions and changed 18 hidden test-case results. Every rule applied equally to both tools, and raw sensitivity results still favored Codex.
  • The unpublished repositories reduce benchmark contamination, but 18 tasks cannot represent every codebase, language, or software engineering workload.
  • Comparable cost data was not available from both subscription CLIs, so no cost comparison is reported.
  • Independent blinded human review of subjective code-quality dimensions is not complete. This page reports objective test, completion, speed, and reliability outcomes only.

Final Recommendation

Choose Codex first if your priority is autonomous completion, hidden-test performance, and speed under a controlled CLI workflow like this one. It won the primary outcome by 25.9 percentage points and the task-level record by 11 wins to 1. Claude Code remains worth testing if your team values its terminal-centered workflow, hooks, permissions, and Anthropic ecosystem. Run both on your own repository before standardizing, because this result covers one controlled configuration rather than every native feature or workload.

⚡ Quick Verdict

Codex won this 108-run controlled CLI benchmark, completing 47 of 54 runs autonomously against Claude Code at 33 of 54. It also finished about 4.3 times faster by median completion time.

Codex achieved an 87.0 percent autonomous success rate against Claude Code at 61.1 percent across 54 runs each. The 25.9-point difference remained positive in a task-clustered 95 percent confidence interval from 11.1 to 42.6 points. Codex won 11 tasks, Claude Code won 1, and 6 were ties. This is a controlled CLI result for the tested versions and configuration, not a claim about every product surface or workload.

Want the full picture? See how Codex and Claude Code stack up against every other tool in Best AI Coding Tools in 2026: A Complete Guide.

🎯 Key Difference

In this controlled CLI setup, the largest practical differences were completion reliability and speed. Codex returned a clean autonomous success more often and had no provider or CLI errors, while Claude Code had 10 such errors and a much longer median completion time.

Feature-by-Feature Comparison

FeatureCodexClaude Code
Autonomous task success✓ Winner47 of 54 runs, 87.0%33 of 54 runs, 61.1%
Adjudicated hidden test cases✓ Winner148 of 156, 94.9%135 of 156, 86.5%
Raw hidden test cases✓ Winner140 of 156, 89.7%125 of 156, 80.1%
Median completion time✓ Winner83.5 seconds357.6 seconds
Provider or CLI errors✓ Winner0 of 54 runs10 of 54 runs
Public regression suites54 of 54 passed54 of 54 passed
Human interventions00
Task-level record✓ Winner11 wins, 1 loss, 6 ties1 win, 11 losses, 6 ties

Head-to-Head Ratings

Scored 1–5 across key dimensions

CodexClaude Code
Autonomous Success
Codex
4.4
Claude Code
3.1
Codex: 47 of 54 runs completed autonomouslyClaude Code: 33 of 54 runs completed autonomously
Hidden Evaluation
Codex
4.7
Claude Code
4.3
Codex: 94.9% adjudicated hidden test-case pass rateClaude Code: 86.5% adjudicated hidden test-case pass rate
Completion Speed
Codex
5
Claude Code
1.2
Codex: 83.5-second medianClaude Code: 357.6-second median
CLI Reliability
Codex
5
Claude Code
4.1
Codex: No provider or CLI errors in 54 canonical runsClaude Code: 10 provider or CLI errors in 54 canonical runs
Regression Avoidance
Codex
5
Claude Code
5
Codex: All 54 public suites passedClaude Code: All 54 public suites passed

Pricing Comparison

Comparable first-party cost data was not available from both subscription CLIs, so this benchmark does not name a cost winner. Both products were tested through paid subscriptions, and actual value will depend on plan limits, model selection, repository size, and task length.

Which Tool Wins for Your Use Case?

Highest autonomous success in this controlled task set
🏆 CodexVisit Site

It completed 47 of 54 runs autonomously, compared with 33 for Claude Code.

Faster completion in this environment
🏆 CodexVisit Site

Its median completion time was 83.5 seconds versus 357.6 seconds.

Terminal-first workflow with hooks and detailed project controls
🏆 Claude CodeVisit Site

The product remains strong for teams that prefer its terminal-centered workflow and customization model.

Choosing across every native product surface
🏆 Run a separate evaluation

This benchmark intentionally disabled vendor-specific features and tested the controlled CLI track only.

The Bottom Line

Codex is the clear winner of this controlled CLI benchmark. It succeeded more often, passed more hidden cases, avoided provider or CLI failures, and finished substantially faster. The result stayed in Codex's favor under raw, unadjudicated scoring. Claude Code still produced six behaviorally successful patches during runs that ended with provider or CLI errors, so its implementation ability was stronger than its primary autonomous score alone suggests.

🏆 Visit CodexOr Try Claude Code

Related Comparisons

More head-to-head comparisons involving these tools

GitHub CopilotVSCursor

GitHub Copilot vs Cursor

CursorVSWindsurf

Cursor vs Windsurf

GitHub CopilotVSTabnine

GitHub Copilot vs Tabnine

CursorVSTabnine

Cursor vs Tabnine

You Might Also Like

More ways to find the right tool

Use-Case GuideBest AI Coding Tools for TeamsRead The Guide →
About This Comparison

Based on publicly available features, pricing, documentation, and real user feedback. Feature tables and scores reflect side-by-side analysis, not subjective preference.

Scored on our 11-dimension framework (1-5 scale). All data verified as of July 2026. Read our full methodology →