My recovered AI history adds up to 76.85 billion tokens. Honestly, that’s kind of ridiculous. But most of those tokens are the models reading the same context again.
Pricing pages make that easy to miss. An agent working in a codebase keeps coming back to the instructions, files, and conversation. Those repeat reads can be much cheaper than the first read.
I wanted to put an API price on that history and see which models gave me better value.
76.85 billion recorded tokens
One square ≈ 1% of the recorded tokens. Ember is cached input.
Remove the cache reads and the number gets a lot less dramatic. What’s left includes fresh input, cache writes, and output, so it isn’t a count of new code either. Neither number tells me how many tasks actually worked.
This is not my bill. The CSV behind this post applies API rates checked on 8 October 2026 to recorded usage. The priced entries come to an API estimate of $39,363–$40,865, with three entries left unpriced. That helps me understand the workload. It doesn’t say what I paid for subscriptions or mean I spent forty thousand dollars.
The models that mattered
The top nine models account for 96% of the recorded tokens. Those are the ones I want to look at; the smaller entries are still in the source table.
Compare nine models in four views
$0.555 / M tokens ÷ Combined 85.61$0.033 / M tokens ÷ Combined 65.21$0.199 / M tokens ÷ Combined 91.25Missing benchmark scores$0.269 / M tokens ÷ Combined 85.84$1.401 / M tokens ÷ Combined 92.94$0.380 / M tokens ÷ Combined 99.71Missing benchmark scores$0.733 / M tokens ÷ Combined 88.00$0.555 / M tokens ÷ AA 46.97$0.033 / M tokens ÷ AA 37.32$0.199 / M tokens ÷ AA 51.83$0.811 / M tokens ÷ AA 29.10$0.269 / M tokens ÷ AA 47.63$1.401 / M tokens ÷ AA 52.67$0.380 / M tokens ÷ AA 57.62$0.807 / M tokens ÷ AA 31.95$0.733 / M tokens ÷ AA 50.78$0.555 / M tokens ÷ Terminal 39.90$0.033 / M tokens ÷ Terminal 11.62$0.199 / M tokens ÷ Terminal 56.06Missing benchmark scores$0.269 / M tokens ÷ Terminal 43.94$1.401 / M tokens ÷ Terminal 59.09$0.380 / M tokens ÷ Terminal 59.60Missing benchmark scores$0.733 / M tokens ÷ Terminal 48.99$0.555 / M tokens ÷ SciCode 57.06$0.033 / M tokens ÷ SciCode 53.59$0.199 / M tokens ÷ SciCode 54.17Missing benchmark scores$0.269 / M tokens ÷ SciCode 57.64$1.401 / M tokens ÷ SciCode 56.48$0.380 / M tokens ÷ SciCode 66.90Missing benchmark scores$0.733 / M tokens ÷ SciCode 56.37$0.555 / M tokens ÷ HLE 49.49$0.033 / M tokens ÷ HLE 39.48$0.199 / M tokens ÷ HLE 52.92$0.811 / M tokens ÷ HLE 30.12$0.269 / M tokens ÷ HLE 47.91$1.401 / M tokens ÷ HLE 54.68$0.380 / M tokens ÷ HLE 61.35$0.807 / M tokens ÷ HLE 39.94$0.733 / M tokens ÷ HLE 54.87$0.555 / M tokens ÷ CRITPt 32.29$0.033 / M tokens ÷ CRITPt 20.57$0.199 / M tokens ÷ CRITPt 31.71$0.811 / M tokens ÷ CRITPt 4.57$0.269 / M tokens ÷ CRITPt 30.86$1.401 / M tokens ÷ CRITPt 31.71$0.380 / M tokens ÷ CRITPt 31.71$0.807 / M tokens ÷ CRITPt 12.57$0.733 / M tokens ÷ CRITPt 29.14$0.555 / M tokens ÷ LCR 84.00$0.033 / M tokens ÷ LCR 83.67$0.199 / M tokens ÷ LCR 83.00$0.811 / M tokens ÷ LCR 77.33$0.269 / M tokens ÷ LCR 83.67$1.401 / M tokens ÷ LCR 80.67$0.380 / M tokens ÷ LCR 84.67$0.807 / M tokens ÷ LCR 78.00$0.733 / M tokens ÷ LCR 79.33All priced entries, including the smaller models.
- GPT-5.6 Sol
- $23,288.95 · 59.2%
- GPT-6 Astra
- $5,063.01 · 12.9%
- Opus 4.5
- $3,470.39 · 8.8%
- All other priced entries
- $7,540.88 · 19.2%
Dollars per BILLION tokensper combined score point
$0.555 / M tokens ÷ Combined 85.61
$0.033 / M tokens ÷ Combined 65.21
$0.199 / M tokens ÷ Combined 91.25
Missing benchmark scores
$0.269 / M tokens ÷ Combined 85.84
$1.401 / M tokens ÷ Combined 92.94
$0.380 / M tokens ÷ Combined 99.71
Missing benchmark scores
$0.733 / M tokens ÷ Combined 88.00
Luna accounts for 12.9% of all recorded tokens and 0.8% of the low-bound API estimate.
The dollar strip shows the three largest low-bound API estimates and groups all other priced entries together. The list ranks the nine main models. Both use all priced entries to calculate their shares. Three unpriced entries stay in the token total. Token and total-dollar views sort from largest to smallest.
Cost per million = low-bound API-equivalent value ÷ recorded tokens × 1,000,000. Includes the recorded mix of cached input, fresh input, cache writes, and output. Sorted from highest to lowest; these are workload prices, not advertised input rates. Every dot uses the same zero-based dollar scale. Dots mark the low price bound; short lines retain the high bound. Movement compares rankings between views, not changes to model prices.
How the value ranking is calculated
Recorded cost per million tokens ÷ selected score. The display uses BILLION tokensso the numbers are easier to read. VALUE KING marks the lowest low-bound ratio. Combined scores average six measures after setting each measure’s highest score in this group to 100. All six scores are required. Individual views use the original benchmark points. Scores and price ranges are in the benchmark section.
Exact counts and API-equivalent values
| Model | Tokens | Cache reads | API value | $/M tokens |
|---|---|---|---|---|
| GPT-5.6 Sol | 41,931,187,448 | 96.9% | $23,288.95–$23,413.02 | $0.56 |
| GPT-5.6 Luna | 9,909,439,234 | 95.0% | $326.50–$326.50 | $0.033 |
| GPT-6.1 Sol | 5,453,682,937 | 96.2% | $1,086.82–$1,086.82 | $0.20 |
| Opus 4.5 | 4,278,733,525 | 94.7% | $3,470.39–$4,293.20 | $0.81 |
| GPT-6 Sol | 3,985,115,186 | 97.3% | $1,071.02–$1,071.02 | $0.27 |
| GPT-6 Astra | 3,614,166,203 | 96.8% | $5,063.01–$5,063.01 | $1.40 |
| Opus 5.5 | 1,797,528,407 | 98.1% | $683.43–$683.43 | $0.38 |
| Opus 4.6 | 1,459,868,074 | 94.9% | $1,177.54–$1,452.43 | $0.81 |
| Opus 5 | 1,339,357,493 | 98.2% | $981.24–$981.24 | $0.73 |
| gpt-5.5 | 1,052,119,878 | 93.2% | $1,016.47–$1,062.44 | $0.97 |
| claude-sonnet-4-5-20250929 | 868,090,992 | 93.2% | $469.54–$599.85 | $0.54 |
| claude-opus-4-7 | 305,100,036 | 93.8% | $326.08–$384.59 | $1.07 |
| gpt-5.6-terra | 304,958,277 | 95.2% | $107.87–$107.87 | $0.35 |
| claude-sonnet-4-6 | 204,837,121 | 94.1% | $136.98–$157.32 | $0.67 |
| claude-opus-4-8 | 106,972,227 | 92.1% | $132.80–$154.35 | $1.24 |
| gpt-5.3-codex-spark | 84,918,505 | 93.9% | Unpriced | — |
| gpt-6-luna | 67,489,853 | 95.8% | $1.08–$1.08 | $0.016 |
| claude-haiku-4-5-20251001 | 57,826,075 | 91.1% | $12.33–$16.03 | $0.21 |
| gpt-5.4 | 17,072,181 | 87.9% | $10.57–$10.57 | $0.62 |
| codex-prior-model-unknown | 10,065,264 | 89.9% | Unpriced | — |
| codex-auto-review | 5,868,292 | 93.7% | Unpriced | — |
| claude-sonnet-5 | 248,170 | 0.0% | $0.61–$0.61 | $2.46 |
GPT-5.6 Sol is the obvious one: 41.93 billion tokens, or 54.6% of the total. GPT-5.6 Luna comes next, with 9.91 billion. Together they account for about two thirds of the recovered history.
Luna accounts for 12.9% of the tokens and only 0.8% of the priced value at the low bound. Counting tokens without looking at which model served them doesn’t tell me much about the cost.
Across the whole dataset, 96.4% of recorded tokens are cached input. I used all tokens, including output, as the denominator.
There is a coverage catch. The recovered Claude history starts on 4 January 2026, has a gap from 6 to 16 June, and runs through 7 October. The recovered Codex history only starts on 6 July. Deleted history may be missing too. So this is a picture of what survived, not a clean comparison of how much I used each provider over the same nine months.
The CSV is also aggregated by model. I can’t turn it into an honest daily usage chart, because the dates of individual requests aren’t there.
Comparing model prices and benchmark scores
The interesting question for me is whether I’m getting more capable models for less money. A cheaper model that needs three attempts and still gets the task wrong isn’t necessarily cheaper in the way I care about.
For a rough comparison, I used six measures from Artificial Analysis: its Intelligence Index, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, CRITPt, and long-context reasoning. They cover terminal work, scientific coding, expert questions, research problems, and reasoning over long inputs. I want to see more than one score when comparing models.
The default view combines them. I set the highest score on each measure among these nine models to 100, then average the six normalized scores with equal weight. A model needs all six scores to enter that comparison. You can switch to any individual benchmark in both the value ranking above and the release chart below.
For each view, I divide API-equivalent dollars per million recorded tokens by the score. The release charts then set the first available model in each family to 100. The ranking uses the price-to-score ratios directly, with the lowest first. The combined score is my own comparison; AA Index already includes several evaluations, so these six measures aren’t independent tests.
I had the “good phones are getting cheap, and cheap phones are getting good” comparison in mind. Switch the chart to Benchmark score to see whether the scores themselves improve, alongside the cost-per-score view.
Model scores and token costs across releases
Compare Sol, Luna, Astra, and Opus. Switch views to see benchmark scores or what each score point costs with my recorded token mix.
Releases are evenly spaced in order, not by elapsed time. Dates are shown below. Vertical ranges in the cost view show the CSV’s low/high price bounds.
Combined · equal weight · cost per score, first release = 100 · lower is better.
Sol
66.3% lower100.0 → 33.7 · low bound
Luna
51.3% lower100.0 → 48.7 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Combined score
- 92.9
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $15.07
Opus
54.2% lower100.0 → 45.8 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Combined score
- 85.61
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Combined score
- 85.84
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 48.3
GPT-6.1 Sol (Max) · released 2026-09-29
- Combined score
- 91.25
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 33.7
GPT-5.6 Luna (Max) · released 2026-07-09
- Combined score
- 65.21
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Combined score
- 65.06
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 48.7
GPT-6 Astra (Max) · released 2026-09-03
- Combined score
- 92.94
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $15.07
Claude Opus 5 (Max) · released 2026-07-24
- Combined score
- 88.00
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 100.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Combined score
- 99.71
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 45.8
Combined · equal weight · benchmark score · higher is better.
Sol
5.6 points higher85.6 → 91.3
Luna
0.1 points lower65.2 → 65.1
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
11.7 points higher88.0 → 99.7
GPT-5.6 Sol (Max) · released 2026-07-09
- Combined score
- 85.61
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Combined score
- 85.84
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 48.3
GPT-6.1 Sol (Max) · released 2026-09-29
- Combined score
- 91.25
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 33.7
GPT-5.6 Luna (Max) · released 2026-07-09
- Combined score
- 65.21
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Combined score
- 65.06
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 48.7
GPT-6 Astra (Max) · released 2026-09-03
- Combined score
- 92.94
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $15.07
Claude Opus 5 (Max) · released 2026-07-24
- Combined score
- 88.00
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 100.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Combined score
- 99.71
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 45.8
AA Intelligence Index v4.3.2 · cost per score, first release = 100 · lower is better.
Sol
67.5% lower100.0 → 32.5 · low bound
Luna
52.5% lower100.0 → 47.5 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Benchmark score
- 52.7
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $26.60
Opus
76.3% lower100.0 → 23.7 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 46.97
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 47.63
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 47.7
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 51.83
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 32.5
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 37.32
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 38.12
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 47.5
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 52.67
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $26.60
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 29.10
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 31.95
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 90.6–111.7
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 50.78
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 51.8
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 57.62
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 23.7
AA Intelligence Index v4.3.2 · benchmark score · higher is better.
Sol
4.9 points higher47.0 → 51.8
Luna
0.8 points higher37.3 → 38.1
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
28.5 points higher29.1 → 57.6
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 46.97
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 47.63
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 47.7
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 51.83
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 32.5
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 37.32
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 38.12
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 47.5
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 52.67
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $26.60
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 29.10
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 31.95
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 90.6–111.7
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 50.78
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 51.8
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 57.62
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 23.7
Terminal-Bench 4.0 · cost per score, first release = 100 · lower is better.
Sol
74.5% lower100.0 → 25.5 · low bound
Luna
55.3% lower100.0 → 44.7 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Benchmark score
- 59.1
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $23.71
Opus
57.3% lower100.0 → 42.7 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 39.90
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 43.94
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 43.9
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 56.06
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 25.5
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 11.62
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 12.63
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 44.7
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 59.09
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $23.71
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 48.99
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 100.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 59.60
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 42.7
Terminal-Bench 4.0 · benchmark score · higher is better.
Sol
16.2 points higher39.9 → 56.1
Luna
1.0 points higher11.6 → 12.6
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
10.6 points higher49.0 → 59.6
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 39.90
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 43.94
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 43.9
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 56.06
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 25.5
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 11.62
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 12.63
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 44.7
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 59.09
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $23.71
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 48.99
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 100.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 59.60
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 42.7
SciCode · cost per score, first release = 100 · lower is better.
Sol
62.2% lower100.0 → 37.8 · low bound
Luna
52.4% lower100.0 → 47.6 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Benchmark score
- 56.5
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $24.80
Opus
56.3% lower100.0 → 43.7 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 57.06
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 57.64
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 47.9
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 54.17
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 37.8
Its SciCode score fell from 57.64 to 54.17, but the lower token price still improves the price-to-score ratio.
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 53.59
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 54.63
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 47.6
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 56.48
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $24.80
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 56.37
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 100.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 66.90
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 43.7
SciCode · benchmark score · higher is better.
Sol
2.9 points lower57.1 → 54.2
Luna
1.0 points higher53.6 → 54.6
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
10.5 points higher56.4 → 66.9
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 57.06
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 57.64
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 47.9
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 54.17
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 37.8
Its SciCode score fell from 57.64 to 54.17, but the lower token price still improves the price-to-score ratio.
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 53.59
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 54.63
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 47.6
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 56.48
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $24.80
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 56.37
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 100.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 66.90
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 43.7
Humanity’s Last Exam · cost per score, first release = 100 · lower is better.
Sol
66.4% lower100.0 → 33.6 · low bound
Luna
50.2% lower100.0 → 49.8 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Benchmark score
- 54.7
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $25.62
Opus
77.0% lower100.0 → 23.0 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 49.49
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 47.91
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 50.0
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 52.92
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 33.6
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 39.48
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 38.51
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 49.8
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 54.68
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $25.62
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 30.12
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 39.94
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 75.0–92.5
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 54.87
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 49.6
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 61.35
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 23.0
Humanity’s Last Exam · benchmark score · higher is better.
Sol
3.4 points higher49.5 → 52.9
Luna
1.0 points lower39.5 → 38.5
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
31.2 points higher30.1 → 61.4
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 49.49
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 47.91
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 50.0
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 52.92
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 33.6
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 39.48
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 38.51
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 49.8
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 54.68
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $25.62
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 30.12
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 39.94
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 75.0–92.5
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 54.87
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 49.6
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 61.35
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 23.0
CRITPt · cost per score, first release = 100 · lower is better.
Sol
63.5% lower100.0 → 36.5 · low bound
Luna
48.6% lower100.0 → 51.4 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Benchmark score
- 31.7
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $44.17
Opus
93.2% lower100.0 → 6.8 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 32.29
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 30.86
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 50.6
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 31.71
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 36.5
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 20.57
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 19.43
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 51.4
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 31.71
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $44.17
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 4.57
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 12.57
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 36.2–44.6
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 29.14
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 14.2
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 31.71
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 6.8
CRITPt · benchmark score · higher is better.
Sol
0.6 points lower32.3 → 31.7
Luna
1.1 points lower20.6 → 19.4
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
27.1 points higher4.6 → 31.7
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 32.29
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 30.86
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 50.6
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 31.71
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 36.5
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 20.57
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 19.43
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 51.4
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 31.71
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $44.17
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 4.57
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 12.57
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 36.2–44.6
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 29.14
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 14.2
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 31.71
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 6.8
Long-context reasoning · cost per score, first release = 100 · lower is better.
Sol
63.7% lower100.0 → 36.3 · low bound
Luna
51.3% lower100.0 → 48.7 · low bound
Astra
One recorded release6 Astra · 3 Sept 2026
- Benchmark score
- 80.7
- Recorded $/million tokens
- $1.40
- $/BILLION tokens / score point
- $17.37
Opus
57.2% lower100.0 → 42.8 · low bound
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 84.00
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 83.67
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 48.6
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 83.00
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 36.3
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 83.67
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 83.33
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 48.7
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 80.67
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $17.37
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 77.33
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 78.00
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 98.6–121.6
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 79.33
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 88.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 84.67
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 42.8
Long-context reasoning · benchmark score · higher is better.
Sol
1.0 points lower84.0 → 83.0
Luna
0.3 points lower83.7 → 83.3
Astra
One recorded release6 Astra · 3 Sept 2026
Opus
7.3 points higher77.3 → 84.7
GPT-5.6 Sol (Max) · released 2026-07-09
- Benchmark score
- 84.00
- Recorded $/M tokens
- $0.56
- Normalized cost per score
- 100.0–100.5
GPT-6 Sol (Max) · released 2026-09-22
- Benchmark score
- 83.67
- Recorded $/M tokens
- $0.27
- Normalized cost per score
- 48.6
GPT-6.1 Sol (Max) · released 2026-09-29
- Benchmark score
- 83.00
- Recorded $/M tokens
- $0.20
- Normalized cost per score
- 36.3
GPT-5.6 Luna (Max) · released 2026-07-09
- Benchmark score
- 83.67
- Recorded $/M tokens
- $0.033
- Normalized cost per score
- 100.0
GPT-6 Luna (Max) · released 2026-09-22
- Benchmark score
- 83.33
- Recorded $/M tokens
- $0.016
- Normalized cost per score
- 48.7
GPT-6 Astra (Max) · released 2026-09-03
- Benchmark score
- 80.67
- Recorded $/M tokens
- $1.40
- $/BILLION tokens / score point
- $17.37
Claude Opus 4.5 (Reasoning) · released 2025-11-24
- Benchmark score
- 77.33
- Recorded $/M tokens
- $0.81–$1.00
- Normalized cost per score
- 100.0–123.7
Claude Opus 4.6 (Max) · released 2026-02-05
- Benchmark score
- 78.00
- Recorded $/M tokens
- $0.81–$0.99
- Normalized cost per score
- 98.6–121.6
Claude Opus 5 (Max) · released 2026-07-24
- Benchmark score
- 79.33
- Recorded $/M tokens
- $0.73
- Normalized cost per score
- 88.0
Claude Opus 5.5 (Max, Default Fallback) · released 2026-09-22
- Benchmark score
- 84.67
- Recorded $/M tokens
- $0.38
- Normalized cost per score
- 42.8
Benchmark scores, model settings, and calculation
Scores retrieved from Artificial Analysis on 8 October 2026. The combined view gives six measures equal weight: AA Index v4.3.2, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, CRITPt, and long-context reasoning. Each measure’s highest score among the nine main models is set to 100, then those six normalized scores are averaged. All six scores are required. AA Index already combines evaluations; this is my own comparison. GPT-6 Luna adds a second Luna release without changing that nine-model reference.
Ratio = API-equivalent dollars per million recorded tokens ÷ benchmark score. Within each family and benchmark, divide by the first available low-bound ratio and multiply by 100. Whiskers retain the CSV’s price range. The score view uses absolute scores on a shared zero-based scale. Lines join evenly spaced releases.
Opus 4.5 and 4.6 have no Terminal-Bench 4.0 or SciCode score in this snapshot. Their combined scores are also unavailable. Those three Opus lines start at Opus 5. The other comparisons start at Opus 4.5. Astra has one recorded release, so its cost view shows absolute values and its score view has a single point.
| Evaluated model setting | AA | Terminal % | SciCode % | HLE % | CRITPt % | LCR % |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol (Max) | 46.97 | 39.90 | 57.06 | 49.49 | 32.29 | 84.00 |
| GPT-5.6 Luna (Max) | 37.32 | 11.62 | 53.59 | 39.48 | 20.57 | 83.67 |
| GPT-6.1 Sol (Max) | 51.83 | 56.06 | 54.17 | 52.92 | 31.71 | 83.00 |
| Claude Opus 4.5 (Reasoning) | 29.10 | — | — | 30.12 | 4.57 | 77.33 |
| GPT-6 Sol (Max) | 47.63 | 43.94 | 57.64 | 47.91 | 30.86 | 83.67 |
| GPT-6 Astra (Max) | 52.67 | 59.09 | 56.48 | 54.68 | 31.71 | 80.67 |
| Claude Opus 5.5 (Max, Default Fallback) | 57.62 | 59.60 | 66.90 | 61.35 | 31.71 | 84.67 |
| Claude Opus 4.6 (Max) | 31.95 | — | — | 39.94 | 12.57 | 78.00 |
| Claude Opus 5 (Max) | 50.78 | 48.99 | 56.37 | 54.87 | 29.14 | 79.33 |
| GPT-6 Luna (Max) | 38.12 | 12.63 | 54.63 | 38.51 | 19.43 | 83.33 |
For the Sol models, the effective token price goes from roughly $0.56 per million for GPT-5.6 Sol to $0.27 for GPT-6 Sol and $0.20 for GPT-6.1 Sol. Their combined scores are 85.6, 85.8, and 91.3. The price-to-score ratio falls by about 66.3% from the first to the last, or 67.5% using AA Index alone.
Sol scores across the benchmarks
| Measure | 5.6 Sol | 6 Sol | 6.1 Sol |
|---|---|---|---|
| Combined | 85.6 | 85.8 | 91.3 |
| AA | 47.0 | 47.6 | 51.8 |
| Terminal | 39.9% | 43.9% | 56.1% |
| SciCode | 57.1% | 57.6% | 54.2% |
| HLE | 49.5% | 47.9% | 52.9% |
| CRITPt | 32.3% | 30.9% | 31.7% |
| LCR | 84.0% | 83.7% | 83.0% |
For Opus, the combined comparison starts at Opus 5, because Opus 4.5 and 4.6 are missing Terminal-Bench and SciCode scores. From Opus 5 to Opus 5.5, the low-bound ratio falls by 54.2%. On AA Index alone, I can compare Opus 4.5 with Opus 5.5, and the drop is 76.3%. The price ranges in the chart matter here: some older Claude records have cache-write counts without a known cache lifetime, so there isn’t one exact price to put on them.
From GPT-5.6 Luna to GPT-6 Luna, AA Index rises from 37.3 to 38.1. Terminal-Bench and SciCode also improve, while HLE, CRITPt, and long-context reasoning fall slightly. The combined score goes from 65.2 to 65.1, so I wouldn’t call that a clear improvement across the board.
Luna’s recorded token cost falls from about $0.033 to $0.016 per million, which brings cost per combined score down 51.3%. Those prices use each model’s own workload: GPT-6 Luna has 67.5 million recorded tokens here, compared with 9.91 billion for GPT-5.6 Luna.
GPT-6 Astra has an AA Index score of about 52.7, close to GPT-6.1 Sol’s 51.8, but its effective token price in this history is about seven times higher. The chart includes its score and cost. There is only one Astra release in the CSV, so I can’t show a trend for that family.
GPT-6.1 Sol’s SciCode score is slightly below GPT-6 Sol’s. Its token price falls enough that the price-to-score ratio still improves.
I used the same version of the AA Index, v4.3.2, for all the models. The evaluated settings are listed under the graph, but my CSV doesn’t record the reasoning setting for each request. Each model also has its own token mix. I’m pairing benchmark scores with observed workloads, rather than testing every model on the same job.
And benchmark points aren’t units of intelligence. A model scoring 50 isn’t “twice as intelligent” as one scoring 25. The ratio helps me think about direction. It doesn’t tell me how much a correct, reviewed pull request costs.
The price of outsourcing the work
If I hand an agent a task in a repository, it may read the instructions, files, and conversation many times. Those reads still count as tokens. But when they qualify as cache hits, charging all of them at the fresh-input price gives me a pretty misleading idea of what the work costs.
GPT-6.1 Sol is a good example. Its base rates in this snapshot are $2 per million input tokens and $10 per million output tokens. If I keep my recorded input/output proportions but price all input as uncached, that works out to $2.03 per million total tokens. With the recorded cache mix, it is $0.20, about 90% lower.
Try taking the cache away below. The model and output share stay the same. A $100 API budget buys roughly 50 million total tokens without cache, or 502 million with my recorded mix.
What if the cache disappeared?
Price one billion total tokens with GPT-6.1 Sol. Keep the output share fixed and change how much input comes from cache.
$0.20 per million total tokens. A $100 API budget buys roughly 502M tokens at this mix.
Rates and calculation
A pricing scenario, not a completed-task estimate. Base rates: $2/M fresh input, $0.10/M cached input, $10/M output. Output stays at 0.3% of total tokens. “My recorded mix” uses the exact input cache-read share, 96.6%; the slider’s manual steps are 0.1 percentage points. Cache writes, long-context premiums, tool fees, and other uplifts are excluded. Cache eligibility is not guaranteed by the slider.
The final graph applies the same uncached comparison to the nine main models. Their API estimates come out roughly 80–91% below the uncached prices, depending on the model and the price bound.
Cheaper attempts and reviews are useful. I still have to check the work, as I wrote about in my post on layers of code review. Cheap tokens don’t fix a bad task description or make an incorrect change useful.
The API estimates here include reported cache writes and recoverable long-context premiums. They leave out missing Codex cache-write uplift, speed premiums, residency uplifts, and tool fees. The comparison also doesn’t measure retries, human review time, or completed tasks. So I wouldn’t call the gap below my savings. It shows how much the token mix changes the price of the same recorded workload.
When I think about outsourcing work to a model, I want to price the workload I actually send it.
Prices with and without cache
Dollars per MILLION total tokens. Both prices keep each model’s recorded input/output proportions.
Every row uses the same dollar scale, from $0 to $12.00. The gap joins two modeled prices for the same token mix.
- Uncached
- $4.04
- Recorded mix
- $0.56
- Uncached
- $0.20
- Recorded mix
- $0.033
- Uncached
- $2.03
- Recorded mix
- $0.20
- Uncached
- $5.01
- Recorded mix
- $0.81–$1.00
- Uncached
- $2.02
- Recorded mix
- $0.27
- Uncached
- $10.11
- Recorded mix
- $1.40
- Uncached
- $4.06
- Recorded mix
- $0.38
- Uncached
- $5.01
- Recorded mix
- $0.81–$0.99
- Uncached
- $5.08
- Recorded mix
- $0.73
Prices and calculation
The uncached comparison uses base input and output list prices from the CSV. The recorded mix uses its API-value low/high bounds, including reported cache writes and recoverable long-context premiums. Ember dots mark the low bound; the heavier ember marks extend to the high bound. Missing Codex cache-write uplift, speed premiums, residency uplifts, and tool fees are excluded. These are modeled API prices, not subscription charges or measured savings against a real uncached run.

