They are what neolabs like Anthropic and OpenAI need you to believe so they can make back their investors’ money.
But You Just Paid the Invoice
Yes, I hear you. You just paid a Claude / Codex / Grok invoice and it was probably in the hundreds or thousands of dollars. I’m not saying those tokens you just paid don’t exist. I’m saying they are not based on a real market mechanism. They’re an expectation, an aspiration, they’re what neolabs hope to receive. How is that even possible?
I tried to explains everything in the video above.
Here’s a short summary, so you know what to expect.
Why the Token Price Even Exists
A SOTA model is made of 2 things: data and compute. Both are extremely expensive. Neolabs took a lot of money from investors to train those models, and they came up with something plausible, and lately, something that can even produce production level code.
But this is horrendously expensive. Now, as a normal user, you wouldn’t pay the real price, simply because it’s prohibitive and it’s still more effective to hire a developer. So neolabs started to sell subscriptions, which have a number of tokens included. It’s like tasting the product. If you want more, you get the price per token.
The best way to understand this is to think at a Ferrari: it’s an extremely expensive car, you cannot afford it probably. So you do not buy the car. You rent ten hours. Renting is the price by the token. The gap between those two prices is unusually large because in the case of AI the product is new and demand is still low.
We Have 3 Layers of Price
- Real cost — billions. Nobody can afford this.
- Lab token price — close to what they need to look solvent in front of investors, not necessarily what the work is really worth.
- Open source / local — often 10× cheaper, 95–98% good enough for coding, analysis, and admin.
The local, private AI is starting to make sense financially, too. Six months ago a local DeepSeek-class box was $25–50k. Now ~$10k (Mac Studio or two DGX Sparks) gets you Qwen / GLM territory near last-gen Opus / GPT.
Something Changed in the Last 6 Months
Now we know what are the real use cases for LLMs: coding, data work, email/meeting cleanup. These are impressive but in and by themselves do not automatically support the rates neolabs are asking for frontier models.
The reality will probably kick in in the next 3–6 months:
- Either neoabs cut prices to defend their market share, or
- Open-source (China + US) takes over the market entirely
The IPO is the key moment and we are a couple of months away from this.
Top comments (2)
Your observation that AI API token prices are aspirational rather than reflecting the actual cost structure really underscores the gap between marketed rates and the real compute expense behind services like Anthropic and OpenAI. Could you elaborate on how pricing models might be adjusted to more transparently represent the true compute and latency costs per token?
At the current stage, LLMs are economically inefficient. The production cost is huge and the use cases that will justify such prices are not yet at there. It may be used in advanced research, if there is enough funding for the said tasks, but not for every day admin tasks. The token price is an attempt to bridge this inequality.
For the token price to go down, I guess we need more adoption in the sense that AI usage will actually create widespread retail value on way more verticals, right now it does in very few ares, like coding.