Deprecated: Using null as an array offset is deprecated, use an empty string instead in /home/u876752588/domains/capria.vc/public_html/wp-content/plugins/jet-engine/includes/components/blocks-views/dynamic-content/manager.php on line 113
- Token expense hurts everyone. Individuals and businesses that cannot spend fall behind, and most providers lose money on every user.
- Building more compute will not solve it. Power is the binding constraint, and the financing behind the buildout is murky and unsustainable.
- Most tasks do not need a frontier model. You do not use a hand grenade to kill a housefly; you use a fly swatter. Small, vertical, and edge models are that swatter.
- This shapes how Capria invests. The “how” of building is increasingly commoditized; the durable question is “what” to build, backed by domain expertise and proprietary access to data or customers.
Token expense is a problem for everyone
Start with the business users, small and large. In a recent post I called out the “Token Divide” where I described how the cost of tokens is separating those that can afford to work with AI at full power from those that cannot – a problem similar to the Digital Divide of the 2000’s. To start, take a look at software engineers, where those with the budget to run agentic workflows day and night now vastly outproduce their equally talented peers who hit their token usage caps by mid-day. But this divide is not just about engineers — it’s also about the other end of the economic ladder, where the majority of the world’s population strives to thrive. Consider small textile trader called Sunita in Nagpur, a tier-2 city in the middle of India. For Sunita, the gap between a free chatbot available via her smartphone’s browser and a capable agentic model that requires a subscription is the difference between writing incrementally better emails and having deep analytical capacity she could never afford to hire. The pattern is consistent across the spectrum : organizations whose employees use AI deeply, and therefore generate token bills, are measurably more efficient and are able to go after new opportunities — stepping far ahead of those like Sunita that cannot spend to keep pace.
An area that gets less attention in the mainstream media is that token expense is a huge problem for AI service providers, from the largest “foundation labs” on down. The economics of AI have inverted the business model that made software the most profitable industry of the last several decades. When I ran the Windows business at Microsoft, our margin on an additional sale was close to 99%. The marginal cost was little more than a CD-ROM, and we didn’t even pay for most of them. When SaaS arrived, margins on an additional customer settled somewhere around 80 to 90%. In both cases, once you had built the thing, serving one more customer cost almost nothing.
AI does not currently work this way. At all. Every token requires high end compute in the cloud, which is decidedly not free, so more consumption means more cost rather than less. The price of compute has fallen over the past two years and per-token prices have fallen, but the benefit of those declines has been more than offset by the rising number of users consuming free AI services, and by the latest AI models that use way more tokens than can be covered by the subscriptions users pay for them. The providers are running hard to stay in the same place, at best. Open AI’s recent IPO delay until 2027 tells us that the markets are waking up to the mismatch between actual and projected revenues and very real capital expenses.
Throwing CapEx at the problem is not the answer
The instinct across the industry right now is to solve this on the supply side, by building out capacity and buying more compute. There are two problems with that approach.
The first is power. Building compute means little if you cannot secure the electricity to run it. As Dunning, the chief AI officer of Hudson River Trading, put it on a recent episode of the Odd Lots podcast, the chips are available; the power to run them is not [other than maybe in China]. He described wanting to expand a data center and being told to go negotiate with the grid first. Even at his firm’s scale, which is smaller (tens of megawatts) relative to the hyperscalers (gigawatts), power was the binding constraint. Generating more power is a known problem that people are working on, but it is not a near-term fix. Datacenters in space are not likely to address this anytime soon, if ever. See my post on that, here.
The second problem is that the current buildout is both inefficient and riskier than it looks. A large amount of GPU capacity sits idle. Compute demand for model training is spiky and leasing contracts are complicated – making it hard to use GPUs efficiently. And the financing behind the buildout is incredibly murky, with layered and/or circular arrangements between data center operators, the clients leasing the compute (hyperscalers, model providers, and others), silicon makers (NVIDIA etc.), debt providers, and insurers. These structures are not well understood even by many of the people inside them. Yes, there is significant debt involved.
I will not pretend to know the timing of a breakdown of this house of cards, and anyone who claims to is guessing. But the structure is unsustainable. At some point one of a few things gives: the funding for near-unlimited buildout tightens, or the financial arrangements come under stress, or enough people notice that capital is being drained from every other sector of the economy to feed this one. Any of these forces a correction, and a correction pushes the industry toward efficiency rather than brute expansion. Because adding compute supply is slow and difficult, a likely place for near-term progress is the demand side: serving the business users and consumers who actually pay to access the model needed to deliver the productivity they need.
Bringing down the cost of intelligence
Some adjustment in market structure is inevitable. The capital-intensive approach cannot hold, and the suppliers themselves will move toward efficiency in time, once the euphoria fades and the capital cycle stops running in their favor. In the meantime, there is useful work to be done closer to the user.
You do not need a hand grenade to kill a housefly. You need a fly swatter.
Most of what a business or a person actually needs from AI does not require a frontier model running in a distant data center. It requires the right amount of intelligence, applied to a specific task, at a cost that makes sense. This is why we expect small language models to keep gaining ground, and it is why we at Capria have been believers in vertical AI for some time: solutions based on small and mid-sized models built for targeted use cases, integrated with application layers that fit the workflow they serve.
The other piece is locality. The reason large models consume so much compute comes down to memory, inference, long context windows, and incredibly expensive model training and retraining. Narrow a problem to a single domain and a lot of those requirements come down.
Edge computing, where capable models run on hardware people already own, is an emerging space. We believe this is a version of a fly swatter.
One version of AI at the edge that’s primarily in research labs is pooling local compute for a targeted application, with the cloud used only for orchestration when a harder task calls for it. Projects like Exopoint at how idle devices on a local network might be pooled into something more capable than any one of them alone. This is one of a handful of such projects we expect to see productized in the coming years. Google’s Gemini Nano and Apple Intelligence run locally on a phone for lighter tasks; each reaches to the cloud only for the complex problems. There are enterprises running open-source models on GPUs they operate themselves, applying those models to the specific problems they need solved, with the added benefit of full data privacy, rather than paying frontier model prices for general intelligence they don’t need.
In a similar vein – 5C Network, in our portfolio, started in the teleradiology business and built a database of more than 21 million cases comprising billions of images. It trained an AI model on that data and built a co-pilot for radiologists. Additionally, They worked with Indian device leader BPL on their new Cortex Rads AI machine that combines BPL’s indigenous imaging hardware with 5C’s AI-native radiology engine to bring intelligent analysis directly into X-ray workflows, turning the machines into smart diagnostic hubs. BPL’s X-ray machine can’t pass the law entrance exam or write poetry – they do not need to. That is the fly swatter, a targeted model running close to the work and sized to the problem.
What this means for how Capria invests
There are usually two questions when you build a product: what to build, and how to build it. The “how” is being commoditized quickly, as the tools for building improve and spread. The harder and more durable question is the “what.” We are backing founders who understand this, and who are building for the paradigm shifts we see emerging. We favor companies solving precise problems with targeted solutions developed via deep domain expertise. We look for proprietary access to data or customers that make these solutions differentiated and sustainable.