The Zero-Marginal Cost AI Era Is Breaking SaaS Economics
The intelligence inside most B2B software now costs pennies per million tokens. So why does the subscription still cost $99 a seat?
- 01The cost to achieve GPT-3.5-level AI performance dropped roughly 100x between late 2022 and early 2025.
- 02Traditional B2B SaaS gross margins remain high at 75% to 85% despite falling underlying compute costs.
- 03AI token savings are often reinvested into complex workflows rather than lowering subscription prices.
- 04Buyers are increasingly weighing in-house API builds against expensive per-seat SaaS subscriptions.

The short answer
B2B SaaS is expensive relative to AI APIs because you are not paying for the model. You are paying for the wrapper, the go-to-market machine, and a pricing model built for a world where software was hard to build. The underlying intelligence many SaaS products now run on has gotten astonishingly cheap. Google's Gemini 2.0 Flash prices at roughly $0.10 per million input tokens and $0.40 per million output tokens, with a genuinely free tier for developers experimenting or running light production traffic. Meanwhile the SaaS subscription wrapped around a similar feature has barely moved.
This is not a minor pricing quirk. It is a structural mismatch that is starting to show.
How cheap has intelligence actually gotten
The numbers are not close. Epoch AI tracked what it costs to hit a fixed level of model performance over time and found the price to reach GPT-3.5-level results on the MMLU benchmark fell from $20 per million tokens in late 2022 to about $0.18 per million tokens with Gemini 2.0 Flash in early 2025, a roughly 100x drop in a little over two years.1 Depending on the benchmark, Epoch found the rate of decline for a fixed performance bar ranged from 9x to 900x per year, with the fastest drops concentrated in the most recent stretch.1
Stanford's AI Index puts a similar number on it from a different angle: inference cost for a system performing at GPT-3.5 level fell more than 280-fold between November 2022 and October 2024, while hardware costs dropped about 30% a year and energy efficiency improved roughly 40% a year.2 Gartner expects the trend to continue, forecasting that running inference on a 1-trillion-parameter model will cost providers more than 90% less in 2030 than it did in 2025.3
None of this means AI itself is free to run at scale, and Gartner's own analyst is blunt about who captures the savings: "the customer isn't going to see all of this money," because as token costs fall, demand shifts toward more complex, token-hungry tasks like agentic workflows that can cost five to thirty times more per query than a simple chatbot request.3 But for the kind of summarization, classification, extraction, and drafting that powers most SaaS AI features today, the marginal cost of the intelligence itself is now close to zero.
What SaaS margins actually look like
Meanwhile, traditional SaaS gross margins have sat in a comfortable, well-defended range for years. The median gross margin for public SaaS companies runs around 72% to 78%, with top-quartile performers clearing 80%, and mature, scale-stage companies often land at 75% to 85%.4 That means for every dollar of subscription revenue, 75 to 85 cents drops to gross profit before a company even spends on sales, marketing, or R&D.
AI features are starting to dent that number, but from a very high floor. Bessemer Venture Partners' 2025 State of AI data shows fast-growing AI "Supernova" companies averaging around 25% gross margin early on, some even negative, while steadier "Shooting Star" companies trend closer to 60%.5 Model providers themselves are not immune: OpenAI's gross margin is estimated at roughly 50%, and Anthropic's around 60%, once inference costs are counted against API and subscription revenue.5 Even the infrastructure layer is being squeezed, with Microsoft disclosing that AI buildout is pressuring its cloud gross margin percentage even as it holds around 69%.5
The pattern across the stack: as you move from chips (roughly 70% margin) up through cloud, models, and finally applications, margins mostly compress as you get further from the token and closer to the customer relationship, until you hit classic subscription software, where margins snap back up toward 75-85% almost regardless of how much of the actual work is now done by a $0.10-per-million-token model.5
Where the markup actually sits
| Layer | Typical gross margin |
|---|---|
| Chips (e.g. Nvidia) | ~70% |
| Cloud infrastructure | ~50-70% |
| Frontier AI models (API) | ~50-60% |
| AI-native "Supernova" apps | ~25% (early) |
| Classic B2B SaaS subscription | 72-85% |
Source: Tanay Jaipuria's 2025 gross margin survey and CloudZero's SaaS benchmark report.54
Why the subscription price hasn't followed the token price down
Three things are keeping SaaS pricing sticky even as the compute underneath it gets cheaper.
Seat-based pricing was never really about compute. Traditional SaaS priced per user per month because the cost driver was engineering and support, not marginal compute per request. A $99 seat price reflected years of R&D amortized across a customer base, not the cost of running any single query. That pricing logic does not automatically update just because a component underneath it got cheaper.
Switching costs, not unit economics, hold prices up. Data lock-in, integrations, procurement relationships, and the sheer hassle of migration let vendors keep prices flat even as their input costs fall. This is standard incumbent economics: the price reflects what the market will bear, not the cost to deliver.
Frontier capability keeps eating the savings. As Gartner's Will Sommer put it, falling token costs "will unlock relatively low-value capabilities that will become embedded in existing ecosystems," but also unlock higher-value applications that are "more expensive, not less," because they require more tokens per task.3 Vendors that pass along the savings from cheaper models often reinvest the freed-up margin into more expensive, more agentic capability rather than lower prices. The SaaS bill stays the same. The thing being delivered for that bill changes.
What this means for buyers
The gap between what the intelligence costs and what the subscription costs is exactly the kind of gap that gets arbitraged away over time, and the arbitrage is already visible in a few places. Companies with in-house engineering capacity are increasingly asking a build-versus-buy question they would not have asked three years ago: can we replicate this $40,000-a-year workflow tool with a few hundred dollars a month in API calls and a couple of engineering sprints. For narrow, well-defined internal workflows, the answer is increasingly yes, which is part of why treating internal software as something you own and operate, rather than rent indefinitely, is becoming a live strategy rather than a hypothetical one. Remy's approach to software ownership reflects this shift directly: the tools a team builds on cheap models should belong to that team, not sit inside someone else's per-seat markup forever.
That does not mean every SaaS subscription is doomed. Deep workflow integration, data gravity, compliance certifications, and genuine product depth are real and defensible. But the pure "access to a model with a UI on top" category of software is now competing against a substitute that costs, in raw compute terms, close to nothing. The bill for that gap will come due somewhere. Increasingly, it is buyers who are starting to ask why they are paying it.
SaaS pricing reflects R&D, sales, support, and switching costs built up over years, not the marginal cost of the underlying compute. Frontier model APIs now run as low as $0.10 to $0.40 per million tokens, while SaaS subscriptions built on those same models have not been repriced downward to match.
Yes. Epoch AI found the price to match GPT-3.5-level performance fell roughly 100x between late 2022 and early 2025, and Stanford's AI Index put the drop for GPT-3.5-level inference at over 280-fold in a similar window.
Because demand shifts toward more complex, token-hungry tasks as costs fall. Agentic workflows can use five to thirty times more tokens per task than a simple chatbot query, so total spend can rise even as the per-token price drops.
Not exactly. Their own margins are under real pressure from AI features, with some AI-native app margins running as low as 25%. But classic subscription software still enjoys margins in the 75-85% range that have not moved much despite cheaper underlying compute.
For narrow, well-defined internal workflows where the value is mostly running data through a model and formatting the output, the economics increasingly favor building in-house with cheap API access over renting a full SaaS subscription.
- 1.LLM inference prices have fallen rapidly but unequally across tasks — Epoch AI
- 2.Optimizing Inference Costs: The Complete Guide — Mirantis
- 3.AI inference costs set to plunge: Gartner — CIO Dive
- 4.SaaS Gross Margin Benchmarks: What To Track In 2026 — CloudZero
- 5.The State of AI Gross Margins in 2025 — Tanay's Newsletter



