---
title: "Introducing Reason Agent 1.1"
description: "On GPT-6 Astra, Reason Agent beats Codex on the Coding Agent Index, 65.1 against 62.5, at 22% lower cost per solved task."
url: "https://reasonmachines.com/blog/introducing-reason-agent-1-1"
section: "Blog"
---

# Introducing Reason Agent 1.1

_On GPT-6 Astra, Reason Agent beats Codex on the Coding Agent Index, 65.1 against 62.5, at 22% lower cost per solved task._

By Reason Machines

October 10, 2026

[Reason Agent](/) is the autonomous software engineer built by Reason Machines.

Today we’re introducing Reason Agent 1.1. On GPT-6 Astra, scored on the same tasks by the same graders, it beats Codex on the Coding Agent Index, 65.1 against 62.5, and costs 15–23% less per attempt than Codex on every benchmark. Pi scores 61.4.[2](#ref-2)

The model sets the pass rate; the harness sets the bill. The same model can cost up to five times more in one harness than in another.[3](#ref-3) On GPT-6 Astra ours costs less per attempt than Codex and Pi, and scores higher.

## Frontier quality at a lower cost

On GPT-6 Astra, Reason Agent has the highest Coding Agent Index of the three harnesses and the lowest cost per solved task: 22% below Codex and 10% below Pi. Cost per attempt is the mean over all 303 tasks, as Artificial Analysis reports cost per task; cost per solved task divides it by the Index. Figure 2, after the Discussion, links each score to its Harbor job.[2](#ref-2)

### Quality vs cost pareto frontier

Up and to the right is better

Figure 1. Coding Agent Index, rounded to a whole number, against cost per task, cheaper to the right, so up and to the right is better. The points are a selection of Artificial Analysis Coding Agent Index v1.5 results at their published cost per task, and Codex on GPT-6 Astra · xhigh from its public Harbor jobs. The blue point is Reason on GPT-6 Astra · xhigh, scored from its Harbor job. Every cost is the mean over all 303 tasks, Artificial Analysis's basis. Hover a point for its values; click it for its source.

## Core principles

Reason Agent 1.1 uses fewer tokens than Codex on every benchmark: on GPT-6 Astra, 23% fewer input tokens per attempt (1.92M against 2.48M) and 15% fewer output tokens.[2](#ref-2) Two principles shape the harness; we have not yet tested how much of the saving each accounts for.

**Parallel tool execution** Reason Agent’s codemode tool runs JavaScript that calls the other tools, so the model can write one program that runs them in parallel and returns only what matters. A search, five file reads and a test run can take one model call.

**A minimal set of tools** The model sees a small set of general coding tools and a short prompt. Everything else a coding platform needs, from orchestrating other sessions and subagents to cloud computers, preview URLs and pull requests with verified screenshots and video, sits behind deferred tools that codemode finds on demand. The whole harness is built to maximise cache eligibility: the prompt and tools never change between turns and each request only appends to the last, so every turn is fully cache-eligible up to its new content.

Same task, same tools, same model. Reason Agent runs tools in parallel inside one program and is done in 3 model calls. A call-per-tool harness returns to the model after every tool and needs 7 model calls.

## Discussion

Model labs can afford the luxury of training a model and its harness together, in reinforcement-learning environments built around that harness.[6](#ref-6) The model grows up on the harness’s tools and habits. That is why a lab’s harness plays best on its own models, and why that home advantage does not travel. A harness built for every model gets no such head start, so it has to get out of the way.

We ask as little of the model as possible: general tools every capable model already knows and no workflow to follow. Everything platform-specific sits behind code and is found when the model needs it. Reason Agent 1.1 is built on [Pi](https://pi.dev). The home advantage is also smaller than it looks: in independent tests, another harness scored best in nine of twelve comparisons.[3](#ref-3)

The cache matters more than it sounds. Every turn re-reads everything before it, so on agent workloads cached input, not output, is most of the bill, and a few points of hit rate move it by double digits: going from 95% to 98% cut one modelled bill by about 17%.[4](#ref-4)

Loading results…

Figure 2. GPT-6 Astra · xhigh: same model, same tasks, same graders. Every benchmark on all its tasks. Reason attempts the provider refused 1 + 10 times score 0, as on Artificial Analysis. Best value in each column is marked; each score links to its Harbor job. Small figures are 95% intervals; Reason’s lead compares the same tasks. [[2]](#ref-2)

−34%median cost per task against Codex at the same effort

GPT-6.1 Sol · Reason vs Codex

ReasonCodex (Artificial Analysis)

Figure 3. GPT-6.1 Sol at every reasoning effort: Coding Agent Index against cost per task (log scale), Reason at the efforts it ran and Codex at every effort Artificial Analysis lists. Hover a point for its values; click it for its Artificial Analysis row or Harbor job.

## Evaluation method

We run the three Coding Agent Index benchmarks, three attempts per task. Our [methodology](https://github.com/reason-machines/benchmark/blob/main/docs/METHODOLOGY.md) is based on [Artificial Analysis’s](https://artificialanalysis.ai/methodology/coding-agents-benchmarking), and the runner and scorer are in our [benchmark repository](https://github.com/reason-machines/benchmark); every result links to its Harbor job.[5](#ref-5) Other harnesses in Figures 1 and 3 are Artificial Analysis’s own runs and prices, except Codex on GPT-6 Astra · xhigh in Figure 1, which comes from its public Harbor jobs; in Figure 2 they are public Harbor jobs on the same tasks.

## References

<a id="ref-1"></a>
[1] Artificial Analysis. [Coding Agents](https://artificialanalysis.ai/agents/coding-agents), retrieved October 3, 2026.

<a id="ref-2"></a>
[2] Harbor Hub. Reason’s runs, one job per model and effort with every trial and trajectory: [GPT-6 Astra xhigh](https://hub.harborframework.com/jobs/0ac70bbf-10d8-5f45-8077-c428f10cb1a1) · [GPT-6.1 Sol xhigh](https://hub.harborframework.com/jobs/439dab82-597f-539e-9c98-7dde2ddd3072) · [GPT-6.1 Sol high](https://hub.harborframework.com/jobs/6c11b2fe-0dd9-53a7-9fe4-2dec41705472) · [GPT-6.1 Sol medium](https://hub.harborframework.com/jobs/b4405006-b07e-56bc-a427-e0c76c7bee12). Scores follow our methodology: an attempt the provider refused or the reward-hacking judge failed scores 0, so a job’s Hub score can be higher. Cost and token totals per attempt match the provider’s usage; these runs predate per-step token records, and their trajectories mask secret-like words as [redacted]. Figure 2’s other scores link their own public jobs.

<a id="ref-3"></a>
[3] Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica and Matei Zaharia. HarnessTax: How Much Does the Harness Matter for Coding Agents? [Post](https://harnesstax.github.io/).

<a id="ref-4"></a>
[4] Diogo. (KV) Cache Rules Everything Around Me. Complete Skeptic, September 2026. [Post](https://www.completeskeptic.com/p/kv-cache-rules-everything-around).

<a id="ref-5"></a>
[5] Reason Machines. Reason on the Coding Agent Index: run and reproduce Reason on DeepSWE v1.1, SWE-Atlas QnA and Terminal-Bench 4.0 with Harbor. [Repository](https://github.com/reason-machines/benchmark).

<a id="ref-6"></a>
[6] OpenAI. Introducing upgrades to Codex, September 15, 2025. OpenAI describes GPT‑5‑Codex as a version of GPT‑5 optimized for agentic coding in Codex. [Post](https://openai.com/index/introducing-upgrades-to-codex/).
