September 24, 2026
AI Works. The Economics Do Not.
The one thing one must understand about AI has little to do with the tech.

By Ignacio de Gregorio
15 min read
AI has many virtues, but making money isn't one of them. But that's not because it's unprofitable by design (it's not).
It's the approach you take.
And start AI legal startup Harvey learned this the painful way: going from reasonably good economic shape to massive losses all of a sudden, in what I call the "Harvey experience" that I believe will become all too common.
I get asked all the time how to avoid bad scares with AI. And my answer is always something like what you're about to read.
Oh, Harvey, what a scare that was!
Reported by Bloomberg, one of the best-known AI startups, Harvey, which operates exclusively as an AI-powered legal assistant, had quite a scare.
Harvey's gross margin fell from about 50% at the beginning of 2026 to negative 50% in June, as customers started to use its AI agents more.
In layman's terms, this means the company went from earning about $0.50 gross for every dollar it charged to being in the red even before subtracting operating costs like wages, stock-based compensation, or marketing.
According to Bloomberg, the deterioration started after the March update that increased customer usage of Harvey's agents. In other words, the product's sudden usage growth made it less "profitable," breaking one of the golden rules of software: economies of scale; the more users one has, the more profitable you should be.
This is literally dying from success, unheard of in software.
The legal AI company's co-founder, Gabe Pereyra, later acknowledged the negative margin but said Harvey brought it back above zero within a single quarter, even as usage doubled month over month โ we'll later see how they did it.
To be clear, Harvey wasn't profitable before the reversal; it wasn't. Instead, the very surprising thing is:
How can a software company have such a dramatic and sudden reversal in its finances?
And the answer is what I want to get to: AI's default pricing model, subscriptions, is broken.
Subscriptions just don't work anymore
The first issue is fixed pricing. Harvey charges negotiated subscriptions per user.
In this pricing model, customers are charged a fixed price on a "take-or-pay basis"; they pay even if they don't use the product.
On the flip side, you guarantee no further charges, even if the user uses the product a lot.
Subscriptions are incredibly popular in software products because they convert much better; customers are more willing to pay because they can predict the cost.
The alternative, consumption-based pricing, or pay-as-you-go models, are a much harder pill for customers to swallow because they are afraid of getting a sudden large bill they did not expect.
But beyond sales conversions, subscriptions have traditionally been great for software companies because the cost of serving each new user was usually very low, and user costs were predictable across your customer pool. In short, subscriptions were high-margin, margins improved as the customer base grew, and cost of goods sold (the cost of serving the product) was predictable.
You may have noticed I'm using the past tense here. But why? Well, because with AI, this is no longer true.
AI breaks the software pricing model, and most AI-powered software will end up being a consumption-based product, or pay-as-you-go.
And the reason is marginal costs. Let me explain.
Power users can bankrupt you
In traditional, non-AI software, most actions a user would take are cheap, infrastructure-wise. Locate the relevant records, check permissions, apply business rules, update a few entries.
Nothing fancy.
These are well-defined, narrow compute processes. Thus, not only is serving the average user cheap, but one can confidently predict that new users will be cheap too.
Yes, in traditional software, users can launch more exhaustive processes like an expensive and enormous report, an optimization job, or a badly constructed query that can be quite taxing, but these are edge cases, not the norm.
In practice, this means that, after an initial heavy investment in infrastructure, that infrastructure can easily accommodate new users. It has to scale, sure, but with diminishing unitary costs.
That is, adding new users results in only gradual growth in computer resources. So, each unit of compute is shared across more users, and per-user (marginal) costs are minimal.
In other words, if we define the required compute as 'F + cN'; fixed costs 'F' plus the marginal, per-user cost 'c' times the number of users N'. Because 'c' is minimal, your required compute is mostly driven by fixed costs.
As fixed costs are diluted across all users, every new user further reduces 'F' per user, which means your margins scale with your customer pool.
Beautiful.
The problem is that this doesn't hold for AI workloads. AI has the worst of both worlds, because:
- The average workload is already very compute-demanding
- An overly-motivated user can spend out of their mind, meaning marginal costs are all but negligible
In other words, margins no longer necessarily grow with the customer pool. Let's tackle each point separately.
AI workloads are no joke
AI workloads are very demanding for three reasons:
- Chip intensity is usually way higher
- These chips also happen to be much more expensive
- A single user can generate an abnormal amount of requests with arbitrarily long durations
Chip intensity refers to the idea that AI inference requires "a lot of GPUs (and CPUs if we're discussing agents) per user." The main culprit is that inference requires large amounts of memory.
This could be an entire separate article, but the point is that each chip's average memory capacity is much smaller than what the average workload requires. This means every user could need dozens of GPUs and, in some extreme cases, all to themselves.
Historically, compute workloads were mostly compute-bound, meaning "how many chips I need" is mostly defined by CPU compute power, and it was rare to find a user requiring more than one CPU chip on a standard basis.
This compute-centric paradigm shaped hardware design for decades, and much more money went into making chips more compute-powerful than memory-capable. In other words, progress was defined by the urgent need for more compute, not more memory.
And then AI came along and asked for the opposite.
The need for more 'memory-centric' designs predates AI, but AI has made them mandatory.
I can't get into details today, but AI is essentially a memory-bound paradigm (or memory-centric, as I like to call it).
If you really want to understand why, I recommend you read this piece I wrote a while back in my newsletter.
And when you combine this with a hardware supply environment unequivocally unprepared for sudden growth in memory demand, we not only get memory becoming the most expensive commodity on the planet, but you also have a state of hardware where "this is the best I've got" falls incredibly short of what AI really needs.
To make the picture a little less ugly, users can be pooled into the same GPU resources (called 'batching'), but this places a huge burden on performance. In layman's terms, you can serve many customers in parallel, but every customer you add to the batch impacts the average speed at which the other customers in the batch get responses.
This trade-off between batching users (and thus making much more efficient use of your chips) and user speed, the latter being called interactivity, can be observed in the throughput vs interactivity example curve below for DeepSeek v4 Pro running on an NVIDIA Rubin server.
To increase response speed per user, you have to reduce batch sizes. This dramatically improves user experience but reduces throughput per chip just as much.
In plain English, by improving the individual experience, the provider makes less money per chip because, in AI workloads, revenue is proportional to workload size.
Usually we try to stay above the 100 tokens/second/user mark; otherwise, the experience becomes really bad for users.
In some cases, you can offer "premium services" where users get extremely fast responses, but that rapidly reduces your ability to serve concurrently, massively impacting your margins unless you charge way more for that "fast mode", which, by the way, is exactly what providers do.
Put another way, AI infrastructure providers face a dichotomy. Unlike non-AI software, where deploying "more CPUs per user" doesn't automatically mean a better user experience or higher direct revenues (e.g., improving query response times by 10 ms won't make your customer suddenly willing to pay double the subscription price), this is very true for AI workloads, and low chip intensity per user usually implies a worse experience and thus customer churn.
Moreover, the fact that you need "more chips per user" is worsened by the fact that these chips, the famous GPUs and other accelerators, are also incredibly expensive.
This means providers must not only spread the investment and operating costs of running these chips across as many users as possible, but these costs, especially the former, are also enormous.
As you can see below, for a 1 GW data center, which previously required a $50 billion investment and with more recent estimates seeing this value growing into $100 billion/GW over the next year or so, these data centers cost around $8 billion/year in capital costs and "only" $900 million/year to run.
Needless to say, capital costs make up the majority of a data center's costs, reaching almost 90% and potentially higher with upcoming platforms like NVIDIA's Rubin GPUs, which are double the price of Blackwell.
Therefore, when paying back an investment in a data center serving AI-powered products, the initial bill is already much higher than in traditional data centers.
Careful, doubling capital costs does not mean worse margins. Rubin GPUs offer way more than 2x performance improvements, so the costs/user can actually fall. Topic for another time.
All things considered, you need more chips per user, and each chip costs much more.
Lovely combo!
The third and final issue is that while "behemoth requests" are rare in non-AI software, they are very common in AI. In fact, with enough ambition, a single user can "hijack" your entire multiple-million-dollar server.
If you're enjoying this article, I think you'll really enjoy my newsletter, where I dive deep into the key trends and insights you need to know that you can't find elsewhere.
Join today!
Subscribe | TheWhiteBox by Nacho de Gregorio The newsletter to stay ahead of the curve in AI
Power users are a mighty enemy
To show how much a single user can move the needle, let me do a quick back-of-the-envelope math exercise.
Consider an apparently harmless request from a user talking to Qwen 3.8 2.4T, a popular Chinese large language model (LLM), asking it to process a sequence of around 1,000,000 tokens, or roughly 750,000 words. For example, a user might want to discuss details of the Harry Potter book saga, which is roughly that size.
Even on a slow day, if that's your only user, you still need the entire model ready on the server. This model has 2.4 trillion parameters and an effective size of roughly 5 Terabytes, or 5 trillion bytes.
An NVIDIA B300 GPU has 288 GB of HBM memory, so you need at least 18 accelerators to deploy the model for inference. Factoring in the example sequence, this sequence has a working memory size (also known as 'KV Cache') of 94 additional GB, which luckily still fits inside the 18 GPUs.
Needing almost twenty tightly connected GPUs forces you to use the largest servers NVIDIA offers**: the GB300 NVL72, which costs roughly $ 4 million for 72 GPUs, or roughly $60k/GPU**.
That means that the hardware required to serve that user costs one million dollars (although in reality you're forced to pay for the entire server).
Despite being a single user, the workload is gigantic. Now group several of these motivated users and your hardware needs skyrocket. And fast.
Nonetheless, the cost distributions OpenAI published for its own employees would scare the devil: the median researcher employee spends $650/day (i.e., at least 50% of the researcher employee base spends $650 or more every day), and the top 10% spends at least $7k/day (valued at API prices).
Particularly telling was an OpenAI employee, Peter Steinberger, who claimed he could spend up to $1.3 million a month on AI.
The implication is that unit costs don't improve with scale. Far from it. More users don't mean better economies of scale, and compute grows in unison with users, so any margin improvements come solely from fixed-cost dilution.
For a provider, it's impossible to know whether the new user is low-usage or will spend more than the previous 1,000. This unpredictability prevented Harvey from forecasting that their gross margins would crash as better models drove higher usage, because what they charged per customer was fixed and thus wouldn't grow with sudden usage growth.
How are you supposed to run a subscription-based business when your gross margins can swing negative out of the blue? Would you invest in that company?
In short, this means compute requirements rise with user growth, and you also start from a higher place (AI chips are more expensive).
Another factor that makes the situation worse is time. In traditional workloads, powerful CPUs process requests and continuously load and offload new ones as they come and go because the workloads aren't very demanding, and CPUs are built to be very fast on a per-request basis.
But with AI, requests can take multiple seconds, sometimes entire minutes or hours, and even full days (see OpenAI's "solution" to Navier-Stokes), so users can effectively hijack not just a large number of resources for themselves, but do so for excruciatingly long times.
In a nutshell, serving AIs is a hot mess. So, next time you wonder why AI is so expensive, feel free to read this article again.
All things considered, trying to serve AI products at fixed prices is like opening an all-you-can-eat buffet for a fixed $20/person but only accepting rugby players and only after they have trained.
Unless you put a limit (meaning it's not an 'all-you-can-eat' buffet), you're sealing your own fate.
Providers do in fact set 'rate limits' meaning the subscription has a fixed-sized resource allocation. But this leads to really bad customer experiences and distrust, worse than simply leading with consumption-based from the very beginning.
There's no way around it. The future of AI pricing models is consumption-based.
In the meantime, there are many things one can do to improve numbers, but the truth is that the industry remains deeply "unserious"; one subsidized by private capital that refuses to evolve into building something sustainable.
Let me explain.
At some point, we need to make money
According to Pereyra, Harvey's President, they improved its margins through a combination of model routing, changes to the software that coordinates agents, and post-training.
Let me translate AI corporate speak**: They improved margins by still charging the same while lowering serving costs**. Only if executives spoke plainly for once, but I digress.
They did so by routing requests to cheaper models, but also by taking what I believe is the only right approach to building a long-lasting AI business: training your own models.
Harvey's August research update offers one example: Tenet, built by further training Moonshot AI's open-weight Kimi K3, scored highly on many critical tasks compared to frontier models.
This is a tale as old as time in AI; no matter how hard frontier lab pundits try to convince you otherwise, depth beats breadth, so training a capable model on the task at hand will outcompete the generalist model that isn't trained on your data.
The problem is that this doesn't solve the larger issue: economics are still broken. They may be less broken than before, but are still broken in principle.
Funnily enough, Harvey's cost of goods sold skyrocketed because they ran OpenAI/Anthropic's models on a consumption-based model; the more you used their models, the more labs charged you, exactly as AI is meant to be charged.
Running a subscription business on consumption-based serving costs, as Harvey and most other application-layer startups insist on doing, works if you only care about closing deals and increasing your top line (revenue) while having zero respect for your bottom line (profits).
This is not a sign of a mature ecosystem.
Many AI startups are living in Wonderland; they aren't actually building a sustainable business, but one that looks good enough in a pitch to a venture capitalist in hopes that they might somehow become sustainable in the future by sheer growth.
When mommy and daddy pay the bills, and you don't have to worry about building a sustainable business, you can get away with cosplaying as a successful entrepreneur, but that doesn't mean you're one because your company is still losing money.
But enough of my rant; I know this is the "Silicon Valley way": grow really fast, and we'll care about profits later.
However, this strategy doesn't work if customer growth doesn't improve margins enough; size isn't the cure. You're just as unprofitable with 10,000 users as you are with 100, just with larger bills.
You do benefit from some margin improvement from fixed costs dilution, but marginal costs are the real problem.
Most concerningly, the industry is reaching a liquidity cliff where VC money is no longer enough and thus debt takes center stage. And when debt investors come in, unlike venture capitalists who mostly care about revenue growth, debt investors care about profits and cash flows, because they depend on those to get their money back.
Therefore, soon, profits will be all that matters, and I'm seriously concerned most AI startups are not only too scared to transition customers to consumption-based models, but are afraid to do so because they know they don't have a sufficiently good moat to push pay-as-you-go pricing (which can easily quadruple costs for the customer).
You're pushing customers from a low-price, predictable environment to higher prices with higher uncertainty in next month's bill. I would be scared too.
So I believe many AI startups today, even billion-dollar ones, are in an "interesting" situation:
They need to transition to consumption-based pricing to make money, but their business is likely no longer attractive for customers if they do.
Is that a good business, my dear reader?
And I insist: routing more of your product to internal models as Harvey did only solves half the problem; your AI, no matter how local it is, still suffers from the exact same illnesses I was discussing above.
I get subscription sexiness, but it's not viable
Look, I get it. I get the urge, as an entrepreneur, to build pricing models that make it easy for customers to sign. A subscription-based model makes deals much easier to pull through.
But one of the key selling points of subscriptions is that, in exchange for capping your revenue per user, they usually benefit from enormous economies of scale, which is why software VCs have always been obsessed with growth no matter what.
The point of "economies of scale" is that fixed costs are diluted across a larger pool of customers, pushing margins up.
But if any individual can spend the equivalent of thousands of others, that solution can't scale predictably and leads to the sudden, massive gross-margin pendulums Harvey experienced.
Subscriptions are attractive to founders for another big reason: investors. They give you and your investors strong recurring revenue guidance, which matters for large start-ups aiming to go public one day or raise the next private funding round, because recurring revenue lets investors discount cash flows well into the future and feel comfortable applying a hefty multiple to your valuation.
I reiterate that I get it.
But as long as subscriptions are the norm, low margins and incredibly high cost spikes with growing usage will be the norm. Unpredictability will be the norm, too.
As you may guess, I'm notoriously skeptical and curious about how Anthropic and OpenAI will handle this very real issue once they go public, especially the latter, which has a much larger subscription-based customer share.
For now, they've indulged in what can't be described any other way than outrageous subscription subsidies while aggressively transitioning larger customers into consumption-based models.
As measured by SemiAnalysis, the equivalent usage to spend your subscription, priced on a consumption basis, can be up to $14,000 for a $200/month ChatGPT subscription.
Anthropic already limits subscriptions to 150 licenses; above that, you're on a pure consumption-based model.
Luckily, they enjoy high margins on the consumption-based (API) products, which in part subsidizes the subscription-based model.
But they're hemorrhaging money nonetheless, proving that scale doesn't solve the problem.
The model is broken. Pls fix.
As the AI industry is still pretty much a private endeavor, and the public side of things is five of the richest companies on the planet spending all their money and the money of others (via debt) to hide the overall unprofitability of it all, this space has gotten away with putting all their effort and interest into showing strong revenue growth.
There's potential for profitability, and some companies are indeed profitable. For example, for companies mostly focused on inference, like Neoclouds, AI can be a very profitable endeavor (e.g., SpaceX's deal with Anthropic) and can generate positive cash**; in cases like CoreWeave or Nebius, they have positive operating cash flows.**
They still aren't "profitable" in the strict sense because they spend more than the cash flow they generate to expand their compute footprint.
But the hard truth is that these companies avoid huge costs like training, data generation, and research, which means the overall "business of AI" still loses money like there's no tomorrow.
The problem is that this industry is selling itself for far less than it costs, thanks to subscription subsidies.
Without subsidies, customers may not be willing to pay more than the AI actually costs, and I say this despite valuing AI products a lot (and paying for them).
But I remain unconvinced that I'm not a minority, and this minority has the gargantuan task of paying for it all. For all these reasons, I tell every client executive I meet: please plan for a full consumption-based model, because that's exactly where you'll end up.
This is not a whim. This is not a threat. It's just plain obvious when you look at how this technology works under the hood.
Prepare for it, or get ready to immerse yourself in your own particular version of the "Harvey experience."
You've been warned.