The Inference Tax
A friend called me yesterday. He'd read my paper on UI testing with Claude Code. It worked beautifully for him, too. He also couldn't use it.
His company gave him a $1,000 API budget. He burned through it in a day. His employer refuses to let him use a personal Max subscription - data policy, IT, the usual. So he sits there knowing the tool exists, knowing it would 10x his work, and he can't have it. The same Anthropic that ships him a transformative product through one door prices him out through the other.
This is not a procurement story. This is the shape of the next decade.
Two prices for the same thing
Every frontier LLM vendor runs two pricing structures. The consumer flat rate - $20-200/month - and the API meter. They are not the same thing. They are not meant to be.
The consumer plan is priced for adoption. Subsidized aggressively, capped softly, designed to win the user and create the muscle memory. The API is priced for enterprise extraction. Calibrated against legal budgets, marketing budgets, and "AI transformation" line items at F500s with money to burn.
Same model. Same compute. Two prices, two markets, calibrated so the vendor's own consumer app is always structurally cheaper than anything you could build on top of their API.
A Claude Code session doing real work burns $50-200 a day at API rates. A Max subscriber doing the same work pays $200 a month. The difference is not your prompt engineering. The difference is that the vendor is not paying retail to themselves.
Who can build on this
A pre-seed startup with two technical founders running modern agentic tooling burns $5,000-15,000 a month on inference before they have a customer, before they have a product, before they have anything. That used to be the entire monthly burn of a two-person company. Now it's the inference budget.
The pitch was: AI lets two founders do what ten used to. The reality: the cost moved from headcount to inference, total burn didn't decrease, and the cost that did move is harder to negotiate, has worse lock-in, and accumulates no institutional knowledge.
Credits programs are a finger in the dam. $5K from Anthropic, $25K from AWS, $10K from Microsoft for Startups. Calibrated to feel generous against legacy SaaS pricing. Totally inadequate against the cost of building something agentic. Two months of runway extension, not a foundation.
And the credits go to exactly the wrong people. To qualify for most of them you need a tier-1 VC on your cap table, an accelerator badge, or a partner intro. The funded startup with $10M in the bank gets $25K of free Anthropic credits. The unfunded founder with $200 in a Mercury account and a real idea gets nothing. The startups that need the credits least are the only ones who get them. The startups that need them most are filtered out by the qualification gate. Welfare for the well-connected. The grassroots era is over, the credentialing era replaced it, and the credentialing gate happens to align almost perfectly with the people who could pay retail anyway.
Meanwhile the funded startup down the street has a special pricing deal on top of the credits. The hyperscaler's portfolio company has internal cost. The incumbent's "AI feature" is paying near-zero marginal cost because they own the model. You, the unfunded founder, are paying retail to compete against everyone whose cost basis is structurally lower.
This is not democratization. This is enclosure.
Cisco or AWS
Two roads, two histories, both lived through.
Cisco in the 1990s built an ecosystem. They sold the boxes, they published the protocols, they certified the partners, they let a thousand integrators and resellers and specialist vendors make money on top of the platform. The platform was valuable because the ecosystem was valuable. Cisco made the routers; everyone else made the careers, the consulting practices, the adjacent products. The pie grew. Cisco took a respectable slice and let everyone else eat.
AWS in the 2010s did the opposite. They built the primitives, watched which startups built profitable products on top, and then launched competing services with names that were one letter off and prices that were structurally lower because AWS doesn't pay AWS. MongoDB became DocumentDB. Elastic became OpenSearch. Redis became ElastiCache. The ecosystem that built AWS got reabsorbed into AWS. The pie did not grow. AWS just claimed dibs.
Cisco's model produced a generation of network engineers and a thousand specialty vendors. AWS's model produced a generation of acquihires and dead companies. Both made the platform owner rich. Only one made the ecosystem rich.
The LLM vendors are at exactly this decision point right now. Their public posture is Cisco - developer days, SDKs, MCP protocols, ecosystem grants, "please build on us." Their pricing is AWS - calibrated so the vendor's own apps always win, so the only third-party developers who survive are the ones too big to crush or too small to notice.
ChatGPT shipping Apps, Atlas, Tasks, Memory. Claude shipping Code, Skills, Projects. Gemini eating every Google surface. Every category that gets traction as a third-party product gets a first-party version within twelve months, priced into the $20 consumer plan, distributed through the vendor's own UI. This is the App Store playbook. Sherlock at scale. The vendors say ecosystem; the pricing says enclosure.
What the pricing tells you
Vendors lie with words. Vendors don't lie with pricing.
If they wanted an ecosystem of independent builders, they'd offer pricing that lets independent builders profit. They'd narrow the consumer-vs-API gap to a normal wholesale-vs-retail spread. They'd offer reserved capacity for startups at near-cost. They'd treat their developer ecosystem as a complement to be cultivated, not a complement to be commoditized.
Instead the gap is widening. Caching helps if you architect for it. Batch tiers help if your workload fits. Volume discounts help if you're already big. None of these help the founder trying to figure out if their idea works.
The pricing tells you what the vendor thinks of you. The vendor thinks of you as a free demo unit while you're hyping their capability, and a margin opportunity once you have revenue, and an acquisition target if you survive past that. None of those is "partner."
The capture is bipartisan
It's tempting to make this an Anthropic problem or an OpenAI problem. It's neither. It's a structural problem with how AI capital is concentrated and how AI economics work.
Three or four labs can train frontier models. Each is entangled with a hyperscaler that provides the compute. Each is under capital pressure to monetize at rates that justify the training cost. Each runs two pricing tiers because the consumer market requires subsidy and the enterprise market tolerates extraction. None of them is the villain individually. The whole structure is the villain.
Regulation will not save you. The "AI safety" frameworks taking shape in DC and Brussels happen to require compliance overhead only large labs can clear. Licensing regimes, export controls, "responsible scaling" - every one of these raises the floor on who can build a frontier model. The incumbents are lobbying for the regulations that will lock them in. This is not conspiracy. This is incentive.
Open models are the only thing keeping the ecosystem honest. Llama, DeepSeek, Qwen, Mistral within twelve months of frontier on most tasks, dramatically cheaper, no per-token rent. If that gap stays small, the rent extraction has a ceiling. If the gap widens - through capital, regulatory moats, or quietly degrading open releases - the rent has no ceiling.
DeepSeek shipped a frontier-adjacent model on a fraction of the training budget the incumbents told us was necessary. That was the most important AI event of the last two years. Not because DeepSeek itself wins. Because the "you need $10B and 100K H100s to compete" story stopped being credible. The incumbents would prefer you forget this. They will tell you the next generation requires twice the capital. Do not believe them without evidence.
What founders should actually do
Stop building in categories the vendor obviously wants. General chat. Coding assistants. Browser agents. Generic "AI for X" plays where X is anything the vendor's consumer app already touches. You're a free demo unit for a vendor product launch.
Build in categories the vendor doesn't want and can't easily enter. Regulated industries. On-prem deployments. Vertical specialties. Data integration into systems the vendor will never touch. Middleware. Plumbing. The boring stuff.
Architect for portability. If swapping vendors is a six-month project, you don't have leverage. If it's a config change, you do. Multi-vendor by default. Open models in the mix from day one. Caching, routing, batching as architectural primitives, not optimizations.
Make your COGS a small fraction of your price. If your inference cost is 30% of revenue, you have a SaaS-with-tax business and the tax goes to a vendor with more power than you. If it's 3%, you have a real business.
Charge for value, not for tokens. The token is the input. The value is the outcome. Customers who pay $50K/year for a measurable outcome don't care that your inference cost $300; customers who pay $20/month for an LLM wrapper will notice every penny.
Assume prices will not drop as fast as the vendors imply. They've been promising 10x drops for two years. The headline drops happen; the quiet capacity gating happens too. Plan as if today's prices are roughly tomorrow's prices, and be pleasantly surprised if they fall.
What breaks the tax
Every rent extraction in tech history ended the same way: somebody found a technical path around it. Not through it, around it. The mainframe rents ended when minicomputers showed up. The minicomputer rents ended when PCs showed up. The on-prem software rents ended when SaaS showed up. The colo rents ended when cloud showed up. The pattern is the same every time: the rent-extractor's pricing is calibrated against the old architecture and breaks when the new architecture arrives. The incumbents never see it coming because they're optimizing the wrong thing.
The inference tax assumes that frontier capability comes from transformer-architecture LLMs trained at hundred-billion-parameter scale on hundred-thousand-GPU clusters. Every assumption in that sentence is contingent. None of it is physics.
Architectures change. Transformers are six years old. They are not the end of the line. State-space models, mixture-of-experts at scales the incumbents aren't running, retrieval-augmented systems that move capability out of weights and into databases, neuro-symbolic hybrids - any one of these could shift the cost curve by an order of magnitude, and several are showing real signs. The big labs have enormous capex sunk into one specific bet. That capex is also a liability if the bet stops being the right one.
Inference hardware is in early innings. Groq's LPUs are doing inference at speeds and costs that the H100 ecosystem can't match for the workloads they fit. Cerebras, SambaNova, Tenstorrent, Etched - each attacking a different angle. NVIDIA's pricing power on training compute is enormous; their pricing power on inference compute is not the same thing, and the cracks are visible. Apple Silicon and the NPU push from Microsoft, Qualcomm, and Intel are quietly moving billions of inference operations off the cloud entirely. The hyperscaler-frontier-lab axis depends on cloud inference staying expensive. It might not.
Open models keep getting better and cheaper to run. The Llama-DeepSeek-Qwen-Mistral cluster is closer to frontier than the incumbents' messaging admits, and the gap has not been widening. DeepSeek shipped a frontier-adjacent model on a fraction of the announced training budget. Llama runs on a Mac. Qwen 3 runs on a laptop. The frontier is a moving target, but the floor under which it can charge rent is also moving - up.
Distillation and specialization undercut the generalist tax. A 7B model fine-tuned for your specific workflow can beat a frontier model at that workflow, while costing 1% as much to run. The generalist frontier model is being asked to do everything for everyone, which is the most expensive way to do anything. Vertical models are the natural response and they're getting easier to produce. A startup that distills a custom small model for their specific use case is structurally cheaper to run than a startup that pipes everything through Claude or GPT.
And the big one: nobody has demonstrated that LLMs are the only path to artificial intelligence, or even the best one. The current consensus is "scale transformers and capability emerges." The consensus has been wrong before. Every prior AI winter happened because the dominant approach hit limits the believers had assured everyone weren't real. We don't know if we're three years from AGI or three years from another winter. The incumbents have priced their products as if the current trajectory continues. That is a bet, not a fact.
Any one of these - new architecture, new hardware, open models maturing, distillation winning, a paradigm shift - could break the inference tax. Two or three of them happening together could destroy it. The vendors know this. It's why they're vertically integrating now, while they can, while the moat is still there. The rush of consumer apps, agent platforms, MCP servers, browser integrations - all of it is racing against the day when the model layer becomes a commodity and the only durable value is in the application layer. Which is to say, in the layer they're currently squeezing.
If you're building outside the moat, time is on your side. Just barely. Not infinitely.
The road
Cisco or AWS. Ecosystem or enclosure. The LLM vendors will choose, and the choice will be made by what their pricing rewards, not what their keynotes promise.
If they choose ecosystem, the next decade is the best startup environment in history. Capital-efficient founders building specialist products in every vertical, with frontier-grade capability available as a utility, growing a pie that the model owners take a respectable slice of.
If they choose enclosure, the next decade is the worst startup environment in history. A handful of well-capitalized partners. The vendor's consumer app eating every category that works. A long tail of founders pricing themselves out before they reach product-market fit. The bootstrapped weirdos who built the last twenty years of software displaced by capital concentration.
I am not betting on the keynotes. I am betting on the pricing.
I am also betting against myself, which is the uncomfortable part. The company I'm building, Yovico, is a simulation environment for your company and its market. Heavily agentic. Burns tokens. The product premise requires inference in the critical path, at volume, across long-running scenarios with state. It is exactly the kind of company the inference tax is designed to break.
So why build it? Because the alternative is to let the only kind of company that survives the inference tax be the kind already inside the vendor moat. Someone has to demonstrate that capital-efficient, vendor-agnostic, agent-heavy products are buildable outside the credentialed ecosystem. Multi-vendor by architecture. Open models for the cheap operations. Frontier models reserved for what genuinely needs them. Routing as a first-class concern, not an optimization. COGS engineered to be a small fraction of price, not the whole P&L.
If I'm wrong about which way the vendors go, the worst case is I built a sensible, portable, capital-efficient product that survives a friendly ecosystem. If I'm right about the vendors but wrong about how long the moat lasts, the worst case is I built early for a market that arrives faster than the rent-extractors can fortify against. If I'm right about both, the worst case is everyone else built on rented land.
The pricing tells you what the vendors plan. The history tells you how it ends. Pay attention to both.