essay

The price tag isn't the cost of AI

Three matching context sheets linked by reuse arrows. The first has a large price tag; the repeated reads have smaller tags.

My recovered AI history adds up to 76.85 billion tokens. Honestly, that’s kind of ridiculous. But most of those tokens are the models reading the same context again.

Pricing pages make that easy to miss. An agent working in a codebase keeps coming back to the instructions, files, and conversation. Those repeat reads can be much cheaper than the first read.

I wanted to put an API price on that history and see which models gave me better value.

76.85 billion recorded tokens

One square ≈ 1% of the recorded tokens. Ember is cached input.

76.85Btokens, including repeated context
96.4% are cache reads. The remaining 2.77B are fresh input, cache writes, and output.

Remove the cache reads and the number gets a lot less dramatic. What’s left includes fresh input, cache writes, and output, so it isn’t a count of new code either. Neither number tells me how many tasks actually worked.

This is not my bill. The CSV behind this post applies API rates checked on 8 October 2026 to recorded usage. The priced entries come to an API estimate of $39,363–$40,865, with three entries left unpriced. That helps me understand the workload. It doesn’t say what I paid for subscriptions or mean I spent forty thousand dollars.

The models that mattered

The top nine models account for 96% of the recorded tokens. Those are the ones I want to look at; the smaller entries are still in the source table.

Compare nine models in four views

GPT-5.6 Sol
41.93B · 54.6%
GPT-5.6 Luna
9.91B · 12.9%
GPT-6.1 Sol
5.45B · 7.1%
Opus 4.5
4.28B · 5.6%
GPT-6 Sol
3.99B · 5.2%
GPT-6 Astra
3.61B · 4.7%
Opus 5.5
1.80B · 2.3%
Opus 4.6
1.46B · 1.9%
Opus 5
1.34B · 1.7%

Luna accounts for 12.9% of all recorded tokens and 0.8% of the low-bound API estimate.

Exact counts and API-equivalent values
ModelTokensCache readsAPI value$/M tokens
GPT-5.6 Sol41,931,187,44896.9%$23,288.95–$23,413.02$0.56
GPT-5.6 Luna9,909,439,23495.0%$326.50–$326.50$0.033
GPT-6.1 Sol5,453,682,93796.2%$1,086.82–$1,086.82$0.20
Opus 4.54,278,733,52594.7%$3,470.39–$4,293.20$0.81
GPT-6 Sol3,985,115,18697.3%$1,071.02–$1,071.02$0.27
GPT-6 Astra3,614,166,20396.8%$5,063.01–$5,063.01$1.40
Opus 5.51,797,528,40798.1%$683.43–$683.43$0.38
Opus 4.61,459,868,07494.9%$1,177.54–$1,452.43$0.81
Opus 51,339,357,49398.2%$981.24–$981.24$0.73
gpt-5.51,052,119,87893.2%$1,016.47–$1,062.44$0.97
claude-sonnet-4-5-20250929868,090,99293.2%$469.54–$599.85$0.54
claude-opus-4-7305,100,03693.8%$326.08–$384.59$1.07
gpt-5.6-terra304,958,27795.2%$107.87–$107.87$0.35
claude-sonnet-4-6204,837,12194.1%$136.98–$157.32$0.67
claude-opus-4-8106,972,22792.1%$132.80–$154.35$1.24
gpt-5.3-codex-spark84,918,50593.9%Unpriced—
gpt-6-luna67,489,85395.8%$1.08–$1.08$0.016
claude-haiku-4-5-2025100157,826,07591.1%$12.33–$16.03$0.21
gpt-5.417,072,18187.9%$10.57–$10.57$0.62
codex-prior-model-unknown10,065,26489.9%Unpriced—
codex-auto-review5,868,29293.7%Unpriced—
claude-sonnet-5248,1700.0%$0.61–$0.61$2.46

GPT-5.6 Sol is the obvious one: 41.93 billion tokens, or 54.6% of the total. GPT-5.6 Luna comes next, with 9.91 billion. Together they account for about two thirds of the recovered history.

Luna accounts for 12.9% of the tokens and only 0.8% of the priced value at the low bound. Counting tokens without looking at which model served them doesn’t tell me much about the cost.

Across the whole dataset, 96.4% of recorded tokens are cached input. I used all tokens, including output, as the denominator.

There is a coverage catch. The recovered Claude history starts on 4 January 2026, has a gap from 6 to 16 June, and runs through 7 October. The recovered Codex history only starts on 6 July. Deleted history may be missing too. So this is a picture of what survived, not a clean comparison of how much I used each provider over the same nine months.

The CSV is also aggregated by model. I can’t turn it into an honest daily usage chart, because the dates of individual requests aren’t there.

Comparing model prices and benchmark scores

The interesting question for me is whether I’m getting more capable models for less money. A cheaper model that needs three attempts and still gets the task wrong isn’t necessarily cheaper in the way I care about.

For a rough comparison, I used six measures from Artificial Analysis: its Intelligence Index, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, CRITPt, and long-context reasoning. They cover terminal work, scientific coding, expert questions, research problems, and reasoning over long inputs. I want to see more than one score when comparing models.

The default view combines them. I set the highest score on each measure among these nine models to 100, then average the six normalized scores with equal weight. A model needs all six scores to enter that comparison. You can switch to any individual benchmark in both the value ranking above and the release chart below.

For each view, I divide API-equivalent dollars per million recorded tokens by the score. The release charts then set the first available model in each family to 100. The ranking uses the price-to-score ratios directly, with the lowest first. The combined score is my own comparison; AA Index already includes several evaluations, so these six measures aren’t independent tests.

I had the “good phones are getting cheap, and cheap phones are getting good” comparison in mind. Switch the chart to Benchmark score to see whether the scores themselves improve, alongside the cost-per-score view.

Model scores and token costs across releases

Compare Sol, Luna, Astra, and Opus. Switch views to see benchmark scores or what each score point costs with my recorded token mix.

Releases are evenly spaced in order, not by elapsed time. Dates are shown below. Vertical ranges in the cost view show the CSV’s low/high price bounds.

Combined · equal weight · cost per score, first release = 100 · lower is better.

Sol

66.3% lower100.0 → 33.7 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

51.3% lower100.0 → 48.7 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Combined score
92.9
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$15.07

Opus

54.2% lower100.0 → 45.8 · low bound

Opus 5
Opus 5.5

Combined · equal weight · benchmark score · higher is better.

Sol

5.6 points higher85.6 → 91.3

5.6 Sol
6 Sol
6.1 Sol

Luna

0.1 points lower65.2 → 65.1

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

11.7 points higher88.0 → 99.7

Opus 5
Opus 5.5

AA Intelligence Index v4.3.2 · cost per score, first release = 100 · lower is better.

Sol

67.5% lower100.0 → 32.5 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

52.5% lower100.0 → 47.5 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Benchmark score
52.7
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$26.60

Opus

76.3% lower100.0 → 23.7 · low bound

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

AA Intelligence Index v4.3.2 · benchmark score · higher is better.

Sol

4.9 points higher47.0 → 51.8

5.6 Sol
6 Sol
6.1 Sol

Luna

0.8 points higher37.3 → 38.1

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

28.5 points higher29.1 → 57.6

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

Terminal-Bench 4.0 · cost per score, first release = 100 · lower is better.

Sol

74.5% lower100.0 → 25.5 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

55.3% lower100.0 → 44.7 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Benchmark score
59.1
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$23.71

Opus

57.3% lower100.0 → 42.7 · low bound

Opus 5
Opus 5.5

Terminal-Bench 4.0 · benchmark score · higher is better.

Sol

16.2 points higher39.9 → 56.1

5.6 Sol
6 Sol
6.1 Sol

Luna

1.0 points higher11.6 → 12.6

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

10.6 points higher49.0 → 59.6

Opus 5
Opus 5.5

SciCode · cost per score, first release = 100 · lower is better.

Sol

62.2% lower100.0 → 37.8 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

52.4% lower100.0 → 47.6 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Benchmark score
56.5
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$24.80

Opus

56.3% lower100.0 → 43.7 · low bound

Opus 5
Opus 5.5

SciCode · benchmark score · higher is better.

Sol

2.9 points lower57.1 → 54.2

5.6 Sol
6 Sol
6.1 Sol

Luna

1.0 points higher53.6 → 54.6

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

10.5 points higher56.4 → 66.9

Opus 5
Opus 5.5

Humanity’s Last Exam · cost per score, first release = 100 · lower is better.

Sol

66.4% lower100.0 → 33.6 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

50.2% lower100.0 → 49.8 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Benchmark score
54.7
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$25.62

Opus

77.0% lower100.0 → 23.0 · low bound

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

Humanity’s Last Exam · benchmark score · higher is better.

Sol

3.4 points higher49.5 → 52.9

5.6 Sol
6 Sol
6.1 Sol

Luna

1.0 points lower39.5 → 38.5

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

31.2 points higher30.1 → 61.4

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

CRITPt · cost per score, first release = 100 · lower is better.

Sol

63.5% lower100.0 → 36.5 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

48.6% lower100.0 → 51.4 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Benchmark score
31.7
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$44.17

Opus

93.2% lower100.0 → 6.8 · low bound

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

CRITPt · benchmark score · higher is better.

Sol

0.6 points lower32.3 → 31.7

5.6 Sol
6 Sol
6.1 Sol

Luna

1.1 points lower20.6 → 19.4

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

27.1 points higher4.6 → 31.7

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

Long-context reasoning · cost per score, first release = 100 · lower is better.

Sol

63.7% lower100.0 → 36.3 · low bound

5.6 Sol
6 Sol
6.1 Sol

Luna

51.3% lower100.0 → 48.7 · low bound

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

Benchmark score
80.7
Recorded $/million tokens
$1.40
$/BILLION tokens / score point
$17.37

Opus

57.2% lower100.0 → 42.8 · low bound

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5

Long-context reasoning · benchmark score · higher is better.

Sol

1.0 points lower84.0 → 83.0

5.6 Sol
6 Sol
6.1 Sol

Luna

0.3 points lower83.7 → 83.3

5.6 Luna
6 Luna

Astra

One recorded release6 Astra · 3 Sept 2026

6 Astra

Opus

7.3 points higher77.3 → 84.7

Opus 4.5
Opus 4.6
Opus 5
Opus 5.5
Benchmark scores, model settings, and calculation

Scores retrieved from Artificial Analysis on 8 October 2026. The combined view gives six measures equal weight: AA Index v4.3.2, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, CRITPt, and long-context reasoning. Each measure’s highest score among the nine main models is set to 100, then those six normalized scores are averaged. All six scores are required. AA Index already combines evaluations; this is my own comparison. GPT-6 Luna adds a second Luna release without changing that nine-model reference.

Ratio = API-equivalent dollars per million recorded tokens ÷ benchmark score. Within each family and benchmark, divide by the first available low-bound ratio and multiply by 100. Whiskers retain the CSV’s price range. The score view uses absolute scores on a shared zero-based scale. Lines join evenly spaced releases.

Opus 4.5 and 4.6 have no Terminal-Bench 4.0 or SciCode score in this snapshot. Their combined scores are also unavailable. Those three Opus lines start at Opus 5. The other comparisons start at Opus 4.5. Astra has one recorded release, so its cost view shows absolute values and its score view has a single point.

Evaluated model settingAATerminal %SciCode %HLE %CRITPt %LCR %
GPT-5.6 Sol (Max)46.9739.9057.0649.4932.2984.00
GPT-5.6 Luna (Max)37.3211.6253.5939.4820.5783.67
GPT-6.1 Sol (Max)51.8356.0654.1752.9231.7183.00
Claude Opus 4.5 (Reasoning)29.10——30.124.5777.33
GPT-6 Sol (Max)47.6343.9457.6447.9130.8683.67
GPT-6 Astra (Max)52.6759.0956.4854.6831.7180.67
Claude Opus 5.5 (Max, Default Fallback)57.6259.6066.9061.3531.7184.67
Claude Opus 4.6 (Max)31.95——39.9412.5778.00
Claude Opus 5 (Max)50.7848.9956.3754.8729.1479.33
GPT-6 Luna (Max)38.1212.6354.6338.5119.4383.33

For the Sol models, the effective token price goes from roughly $0.56 per million for GPT-5.6 Sol to $0.27 for GPT-6 Sol and $0.20 for GPT-6.1 Sol. Their combined scores are 85.6, 85.8, and 91.3. The price-to-score ratio falls by about 66.3% from the first to the last, or 67.5% using AA Index alone.

Sol scores across the benchmarks

Measure5.6 Sol6 Sol6.1 Sol
Combined85.685.891.3
AA47.047.651.8
Terminal39.9%43.9%56.1%
SciCode57.1%57.6%54.2%
HLE49.5%47.9%52.9%
CRITPt32.3%30.9%31.7%
LCR84.0%83.7%83.0%

For Opus, the combined comparison starts at Opus 5, because Opus 4.5 and 4.6 are missing Terminal-Bench and SciCode scores. From Opus 5 to Opus 5.5, the low-bound ratio falls by 54.2%. On AA Index alone, I can compare Opus 4.5 with Opus 5.5, and the drop is 76.3%. The price ranges in the chart matter here: some older Claude records have cache-write counts without a known cache lifetime, so there isn’t one exact price to put on them.

From GPT-5.6 Luna to GPT-6 Luna, AA Index rises from 37.3 to 38.1. Terminal-Bench and SciCode also improve, while HLE, CRITPt, and long-context reasoning fall slightly. The combined score goes from 65.2 to 65.1, so I wouldn’t call that a clear improvement across the board.

Luna’s recorded token cost falls from about $0.033 to $0.016 per million, which brings cost per combined score down 51.3%. Those prices use each model’s own workload: GPT-6 Luna has 67.5 million recorded tokens here, compared with 9.91 billion for GPT-5.6 Luna.

GPT-6 Astra has an AA Index score of about 52.7, close to GPT-6.1 Sol’s 51.8, but its effective token price in this history is about seven times higher. The chart includes its score and cost. There is only one Astra release in the CSV, so I can’t show a trend for that family.

GPT-6.1 Sol’s SciCode score is slightly below GPT-6 Sol’s. Its token price falls enough that the price-to-score ratio still improves.

I used the same version of the AA Index, v4.3.2, for all the models. The evaluated settings are listed under the graph, but my CSV doesn’t record the reasoning setting for each request. Each model also has its own token mix. I’m pairing benchmark scores with observed workloads, rather than testing every model on the same job.

And benchmark points aren’t units of intelligence. A model scoring 50 isn’t “twice as intelligent” as one scoring 25. The ratio helps me think about direction. It doesn’t tell me how much a correct, reviewed pull request costs.

The price of outsourcing the work

If I hand an agent a task in a repository, it may read the instructions, files, and conversation many times. Those reads still count as tokens. But when they qualify as cache hits, charging all of them at the fresh-input price gives me a pretty misleading idea of what the work costs.

GPT-6.1 Sol is a good example. Its base rates in this snapshot are $2 per million input tokens and $10 per million output tokens. If I keep my recorded input/output proportions but price all input as uncached, that works out to $2.03 per million total tokens. With the recorded cache mix, it is $0.20, about 90% lower.

Try taking the cache away below. The model and output share stay the same. A $100 API budget buys roughly 50 million total tokens without cache, or 502 million with my recorded mix.

What if the cache disappeared?

Price one billion total tokens with GPT-6.1 Sol. Keep the output share fixed and change how much input comes from cache.

GPT-6.1 Sol1,000M tokens
All input uncached$2,027.40
Recorded cache mix$199.20

$0.20 per million total tokens. A $100 API budget buys roughly 502M tokens at this mix.

Rates and calculation

A pricing scenario, not a completed-task estimate. Base rates: $2/M fresh input, $0.10/M cached input, $10/M output. Output stays at 0.3% of total tokens. “My recorded mix” uses the exact input cache-read share, 96.6%; the slider’s manual steps are 0.1 percentage points. Cache writes, long-context premiums, tool fees, and other uplifts are excluded. Cache eligibility is not guaranteed by the slider.

The final graph applies the same uncached comparison to the nine main models. Their API estimates come out roughly 80–91% below the uncached prices, depending on the model and the price bound.

Cheaper attempts and reviews are useful. I still have to check the work, as I wrote about in my post on layers of code review. Cheap tokens don’t fix a bad task description or make an incorrect change useful.

The API estimates here include reported cache writes and recoverable long-context premiums. They leave out missing Codex cache-write uplift, speed premiums, residency uplifts, and tool fees. The comparison also doesn’t measure retries, human review time, or completed tasks. So I wouldn’t call the gap below my savings. It shows how much the token mix changes the price of the same recorded workload.

When I think about outsourcing work to a model, I want to price the workload I actually send it.

Prices with and without cache

Dollars per MILLION total tokens. Both prices keep each model’s recorded input/output proportions.

All input uncachedRecorded mix · low/high bounds

Every row uses the same dollar scale, from $0 to $12.00. The gap joins two modeled prices for the same token mix.

GPT-5.6 Sol86% lower
Uncached
$4.04
Recorded mix
$0.56
GPT-5.6 Luna84% lower
Uncached
$0.20
Recorded mix
$0.033
GPT-6.1 Sol90% lower
Uncached
$2.03
Recorded mix
$0.20
Opus 4.580–84% lower
Uncached
$5.01
Recorded mix
$0.81–$1.00
GPT-6 Sol87% lower
Uncached
$2.02
Recorded mix
$0.27
GPT-6 Astra86% lower
Uncached
$10.11
Recorded mix
$1.40
Opus 5.591% lower
Uncached
$4.06
Recorded mix
$0.38
Opus 4.680–84% lower
Uncached
$5.01
Recorded mix
$0.81–$0.99
Opus 586% lower
Uncached
$5.08
Recorded mix
$0.73
Prices and calculation

The uncached comparison uses base input and output list prices from the CSV. The recorded mix uses its API-value low/high bounds, including reported cache writes and recoverable long-context premiums. Ember dots mark the low bound; the heavier ember marks extend to the high bound. Missing Codex cache-write uplift, speed premiums, residency uplifts, and tool fees are excluded. These are modeled API prices, not subscription charges or measured savings against a real uncached run.