Who actually pays for inference?
Free tiers, student credits, and ‘unlimited’ chat are not gifts. They are a bet that someone else — advertisers, enterprises, or tomorrow’s you — will cover the meter.

The industry’s favorite customer is a person who types all day and pays nothing. The industry’s actual customer is whoever can be made to pay for that person’s tokens without noticing the line item.
That is not a metaphor. It is the structure of the last two years of consumer AI: subsidized chat up front, enterprise contracts and ads to catch the bleeding, and a quiet hope that hardware prices fall faster than product managers add features.
A token is not free because a lab is kind. A token is free because the bill has not been assigned yet.
The meter nobody shows you
Inference cost is not one number. It is utilization, memory, routing, cache hit rate, and the length of the reply the model decides you deserved. A “reasoning” SKU that thinks in public can cost several times a fast model that would have been good enough. Product teams discovered that users click the expensive button when it is the default. Finance teams discovered it a sprint later.
Closed APIs hide this behind a price per million tokens that still assumes you know how many tokens a task takes. Most buyers do not. They buy a seat, or a “pro” plan, and treat the model like electricity they cannot read on the meter.
Open weights move the meter onto your own GPUs, which is clarifying and also how people learn that “free” was never free. You pay in capital, power, and the engineer who keeps the batcher from falling over on Monday.
Three ways the bill lands
Enterprises eat it as software. The pitch is productivity. The invoice is seats plus overage plus a professional-services team that teaches the model not to invent policy. This is the healthiest version of the business, which is why every consumer lab is trying to become a Salesforce with a better chatbot.
Advertisers will eat some of it, eventually, in products that can tolerate a recommendation in the margin. That path requires inventory people actually look at. It also requires the user to keep chatting after they notice the inventory.
Users eat it when the subsidy ends. Price increases get dressed up as “higher limits” and “priority access.” The honest version is: the free tier was a customer-acquisition cost, and the acquisition cost went up.
The part that does not show up in the launch blog
There is a fourth payer: the public, if energy markets and data-center siting keep being treated as someone else’s externality. That is not an argument against building. It is an argument for putting the power purchase on the same slide as the token price.
If you want to know whether a lab is a product company or a subsidy, do not read the system card. Read who flinches when utilization spikes on a Sunday. That is who pays for inference.