Nobody Buys Tokens
Two model launches this week. One lab named its new pair after the sun and the moon. The other lab has the word Space in its name. I read both announcements the way I read everything now: find the unit first, then look at the number.
The sun and the moon, then. GPT-6 Sol costs $2 per million input tokens and $10 per million output. Luna costs ten cents and fifty cents. The headline is a 50% price cut, and the exact wording is “50% compared with their GPT‑5.6 promotional pricing.” Half off the promotional price. That is a real discount. It is also a sentence written by someone who has met a CFO.
Grok 4.7 is $2 in and $6 out, under the tagline “Twice as fast, at half the price of comparable models.” Comparable is doing a lot of work in that sentence. It always is.
Here is what neither price table shows you. Every one of these models has a knob. OpenAI calls it reasoning effort. It is one line in a config file, and it is the most expensive line in the file. Turn it up and the model thinks longer before it answers. The thinking comes back as tokens you never get to read, billed at the output rate. You are paying premium rates for the model’s inner monologue, like a therapist who invoices you for their own thoughts.
How much monologue? Artificial Analysis reported that Grok 4.7, at its xHigh setting, used about 81,000 output tokens per Intelligence Index task — substantially more, in its testing, than Grok 4.6 High or GPT-6 Astra Max. At $6 per million, that is roughly 49 cents of output per task, before a single input token is counted. The per-token price is low. The per-task bill is the per-token price multiplied by a number that never makes the headline.
To be fair to the sun and the moon, the most useful sentence in OpenAI’s post is not the price. It is further down, in the benchmark section. On OSWorld 2.0 offline, Sol at xhigh effort “achieves a similar score to Claude Opus 5 at medium effort—60.5% versus 60.3%—at approximately 80% lower cost per task.” Two-tenths of a point. I have seen grown men in a bar argue less about an inch. But look at the method: match the score first, then compare the bill per finished task. That is the only comparison that means anything, and it lives well below the headline.
I learned this the expensive way. I once moved a nightly batch job to a model that cost half as much per token, and the invoice went up. The new model was cheaper per word and chattier per question, and the job did not care about words. It cared about being done.
So this is what I do now, and it takes one afternoon. Twenty real tasks from my own backlog. The effort level I would actually ship, not the one on the chart. Run them all, then divide the invoice by twenty. That number is the price. Everything on the launch page is a brochure for it.
Nobody buys tokens. You buy a finished task and get billed in tokens, at an exchange rate set by a knob whose default you probably never chose. The sun, the moon and the rocket all come with a meter.
Read the meter.
Terminal’s still open.