Codex vs Claude Code: 108-Run Coding Benchmark
We tested both coding agents on the same 18 unpublished TypeScript, Python, and Go tasks, with three clean repetitions per tool and no human steering.
Research curated by Etienne Tawong · Verified July 2026 · Our methodology →
Codex won the pre-registered primary outcome at 87.0 percent autonomous success versus Claude Code at 61.1 percent. The task-clustered 95 percent confidence interval for the difference ran from 11.1 to 42.6 percentage points in Codex's favor.
Directly verified labels mark findings backed by the frozen harness, logs, patches, and evaluator results.
Primary Outcome
| Category | Weight | Codex | Claude Code |
|---|---|---|---|
| Autonomous task success | 100 | 87 | 61.1 |
| Raw total | 100 | 87 | 61.1 |
| Normalized (/100) | 87.0 | 61.1 |
The page score is the preregistered primary outcome: autonomous task success across 54 canonical runs per tool.
How We Tested
- The controlled matrix used 18 unpublished tasks: six TypeScript, six Python, and six Go. Each tool ran every task three times from a clean repository snapshot, producing 108 canonical runs.
- Both tools received identical prompts, public tests, hidden evaluators, caches, machine resources, and wall-clock limits. The schedule was randomized once from the frozen seed 20260719.
- Vendor-specific extensions were disabled. There was no web search, browser access, MCP, plugins, project-specific instructions, subagents, cross-session history, or human steering after prompt submission.
- Codex CLI 0.144.6 used gpt-5.6-sol. Claude Code 2.1.215 used claude-opus-4-8 at maximum effort. Shell network access was blocked except for each product's control plane.
- Autonomous success required a clean completion, a buildable patch, the full public regression suite, and the complete adjudicated hidden suite. This is stricter than the frozen 90 percent hidden-criteria threshold.
- Statistics use paired task evidence. Confidence intervals resample whole tasks so the three repetitions of one task are not treated as independent tasks.
What Each Tool Produced
Screenshots captured during the run. Click any image to view it full size.
Codex completed a TypeScript preferences task with the public and hidden evaluators passing.
Claude Code completed the paired TypeScript preferences task and returned a passing implementation.
Codex completed a Python cursor task under the frozen harness without human intervention.
Claude Code completed the paired Python cursor task under the same harness and limits.
Codex completed a Go configuration task with the public and hidden evaluators passing.
Claude Code completed the paired Go configuration task with the same evaluator setup.
Direct Observations
- Codex completed 47 of 54 runs autonomously, or 87.0 percent. Claude Code completed 33 of 54, or 61.1 percent.
- The success-rate difference was 25.9 percentage points. Its task-clustered 95 percent bootstrap interval was 11.1 to 42.6 points in Codex's favor.
- Codex won 11 tasks, Claude Code won 1, and 6 were ties. The task-level two-sided exact sign-test p-value was 0.006348.
- Codex passed 148 of 156 adjudicated hidden test cases, or 94.9 percent. Claude Code passed 135 of 156, or 86.5 percent.
- Codex's median completion time was 83.5 seconds. Claude Code's was 357.6 seconds, about 4.3 times longer.
- Codex had no provider or CLI errors in 54 canonical runs. Claude Code had 10. Quota exhaustion was retained as an outcome under the frozen protocol.
- Without any adjudications, Codex still led on autonomous success at 72.2 percent versus 53.7 percent.
Functional Results
Each finding below is backed by the frozen run corpus and objective evaluator output.
- Codex: 47 of 54 autonomous successes, 87.0 percent
- Claude Code: 33 of 54 autonomous successes, 61.1 percent
- Difference: 25.9 percentage points for Codex
- Task-clustered 95 percent interval: 11.1 to 42.6 points
- Codex: 148 of 156 adjudicated hidden test cases passed
- Claude Code: 135 of 156 adjudicated hidden test cases passed
- Both tools passed all 54 public regression suites
- Raw hidden results also favored Codex, 89.7 percent to 80.1 percent
- Codex median completion time: 83.5 seconds
- Claude Code median completion time: 357.6 seconds
- Codex provider or CLI errors: 0 of 54
- Claude Code provider or CLI errors: 10 of 54
Limitations Of This Benchmark
- This is the controlled CLI track. Native product features such as cloud tasks, editor integrations, hooks, skills, subagents, and MCP were deliberately disabled and were not scored.
- The result is a July 2026 snapshot of specific tool versions, model identifiers, subscription accounts, and one Apple silicon machine.
- Provider limits affected Claude Code during the run window. They were retained because the protocol preregistered quota exhaustion and tool errors as outcomes.
- Five adjudication rules addressed unstated evaluator assumptions and changed 18 hidden test-case results. Every rule applied equally to both tools, and raw sensitivity results still favored Codex.
- The unpublished repositories reduce benchmark contamination, but 18 tasks cannot represent every codebase, language, or software engineering workload.
- Comparable cost data was not available from both subscription CLIs, so no cost comparison is reported.
- Independent blinded human review of subjective code-quality dimensions is not complete. This page reports objective test, completion, speed, and reliability outcomes only.
Final Recommendation
Choose Codex first if your priority is autonomous completion, hidden-test performance, and speed under a controlled CLI workflow like this one. It won the primary outcome by 25.9 percentage points and the task-level record by 11 wins to 1. Claude Code remains worth testing if your team values its terminal-centered workflow, hooks, permissions, and Anthropic ecosystem. Run both on your own repository before standardizing, because this result covers one controlled configuration rather than every native feature or workload.
⚡ Quick Verdict
Codex won this 108-run controlled CLI benchmark, completing 47 of 54 runs autonomously against Claude Code at 33 of 54. It also finished about 4.3 times faster by median completion time.
Codex achieved an 87.0 percent autonomous success rate against Claude Code at 61.1 percent across 54 runs each. The 25.9-point difference remained positive in a task-clustered 95 percent confidence interval from 11.1 to 42.6 points. Codex won 11 tasks, Claude Code won 1, and 6 were ties. This is a controlled CLI result for the tested versions and configuration, not a claim about every product surface or workload.
Want the full picture? See how Codex and Claude Code stack up against every other tool in Best AI Coding Tools in 2026: A Complete Guide.
In this controlled CLI setup, the largest practical differences were completion reliability and speed. Codex returned a clean autonomous success more often and had no provider or CLI errors, while Claude Code had 10 such errors and a much longer median completion time.
Feature-by-Feature Comparison
Head-to-Head Ratings
Scored 1–5 across key dimensions
Pricing Comparison
Comparable first-party cost data was not available from both subscription CLIs, so this benchmark does not name a cost winner. Both products were tested through paid subscriptions, and actual value will depend on plan limits, model selection, repository size, and task length.
Which Tool Wins for Your Use Case?
It completed 47 of 54 runs autonomously, compared with 33 for Claude Code.
Its median completion time was 83.5 seconds versus 357.6 seconds.
The product remains strong for teams that prefer its terminal-centered workflow and customization model.
This benchmark intentionally disabled vendor-specific features and tested the controlled CLI track only.
The Bottom Line
Codex is the clear winner of this controlled CLI benchmark. It succeeded more often, passed more hidden cases, avoided provider or CLI failures, and finished substantially faster. The result stayed in Codex's favor under raw, unadjudicated scoring. Claude Code still produced six behaviorally successful patches during runs that ended with provider or CLI errors, so its implementation ability was stronger than its primary autonomous score alone suggests.
Related Comparisons
More head-to-head comparisons involving these tools
You Might Also Like
More ways to find the right tool
Based on publicly available features, pricing, documentation, and real user feedback. Feature tables and scores reflect side-by-side analysis, not subjective preference.



