Reason Agent is the autonomous software engineer built by Reason Machines.
Today we’re introducing Reason Agent 1.1. On GPT-6 Astra, scored on the same tasks by the same graders, it beats Codex on the Coding Agent Index, 65.1 against 62.5, and costs 15–23% less per attempt than Codex on every benchmark. Pi scores 61.4.2
The model sets the pass rate; the harness sets the bill. The same model can cost up to five times more in one harness than in another.3 On GPT-6 Astra ours costs less per attempt than Codex and Pi, and scores higher.
Frontier quality at a lower cost
On GPT-6 Astra, Reason Agent has the highest Coding Agent Index of the three harnesses and the lowest cost per solved task: 22% below Codex and 10% below Pi. Cost per attempt is the mean over all 303 tasks, as Artificial Analysis reports cost per task; cost per solved task divides it by the Index. Figure 2, after the Discussion, links each score to its Harbor job.2
Quality vs cost pareto frontier
Up and to the right is betterCore principles
Reason Agent 1.1 uses fewer tokens than Codex on every benchmark: on GPT-6 Astra, 23% fewer input tokens per attempt (1.92M against 2.48M) and 15% fewer output tokens.2 Two principles shape the harness; we have not yet tested how much of the saving each accounts for.
Parallel tool execution Reason Agent’s codemode tool runs JavaScript that calls the other tools, so the model can write one program that runs them in parallel and returns only what matters. A search, five file reads and a test run can take one model call.
A minimal set of tools The model sees a small set of general coding tools and a short prompt. Everything else a coding platform needs, from orchestrating other sessions and subagents to cloud computers, preview URLs and pull requests with verified screenshots and video, sits behind deferred tools that codemode finds on demand. The whole harness is built to maximise cache eligibility: the prompt and tools never change between turns and each request only appends to the last, so every turn is fully cache-eligible up to its new content.
One program runs four reads and searches in parallel, then edits and tests.
The model is called again after every single tool.
Discussion
Model labs can afford the luxury of training a model and its harness together, in reinforcement-learning environments built around that harness.6 The model grows up on the harness’s tools and habits. That is why a lab’s harness plays best on its own models, and why that home advantage does not travel. A harness built for every model gets no such head start, so it has to get out of the way.
We ask as little of the model as possible: general tools every capable model already knows and no workflow to follow. Everything platform-specific sits behind code and is found when the model needs it. Reason Agent 1.1 is built on Pi. The home advantage is also smaller than it looks: in independent tests, another harness scored best in nine of twelve comparisons.3
The cache matters more than it sounds. Every turn re-reads everything before it, so on agent workloads cached input, not output, is most of the bill, and a few points of hit rate move it by double digits: going from 95% to 98% cut one modelled bill by about 17%.4
Loading results…
GPT-6.1 Sol · Reason vs Codex
Evaluation method
We run the three Coding Agent Index benchmarks, three attempts per task. Our methodology is based on Artificial Analysis’s, and the runner and scorer are in our benchmark repository; every result links to its Harbor job.5 Other harnesses in Figures 1 and 3 are Artificial Analysis’s own runs and prices, except Codex on GPT-6 Astra · xhigh in Figure 1, which comes from its public Harbor jobs; in Figure 2 they are public Harbor jobs on the same tasks.
References
- Artificial Analysis. Coding Agents, retrieved October 3, 2026.
- Harbor Hub. Reason’s runs, one job per model and effort with every trial and trajectory: GPT-6 Astra xhigh · GPT-6.1 Sol xhigh · GPT-6.1 Sol high · GPT-6.1 Sol medium. Scores follow our methodology: an attempt the provider refused or the reward-hacking judge failed scores 0, so a job’s Hub score can be higher. Cost and token totals per attempt match the provider’s usage; these runs predate per-step token records, and their trajectories mask secret-like words as [redacted]. Figure 2’s other scores link their own public jobs.
- Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica and Matei Zaharia. HarnessTax: How Much Does the Harness Matter for Coding Agents? Post.
- Diogo. (KV) Cache Rules Everything Around Me. Complete Skeptic, September 2026. Post.
- Reason Machines. Reason on the Coding Agent Index: run and reproduce Reason on DeepSWE v1.1, SWE-Atlas QnA and Terminal-Bench 4.0 with Harbor. Repository.
- OpenAI. Introducing upgrades to Codex, September 15, 2025. OpenAI describes GPT‑5‑Codex as a version of GPT‑5 optimized for agentic coding in Codex. Post.