Most SaaS companies add AI features without understanding what they cost at the customer level. Learn how inference costs, pricing models, and usage patterns are quietly compressing margins — and what finance teams can do about it.
Adding AI features to a SaaS product doesn’t automatically make it more profitable per customer — it makes it more expensive to serve. Vertical SaaS companies that typically ran at 75–82% gross margins in 2023 are now reporting blended margins in the 60–70% range after embedding AI. Multiple public SaaS companies have attributed 6–9 points of gross margin compression directly to AI feature costs.
Traditional SaaS has targeted 70–80% gross margins. AI-native products are tracking toward the low 50s in 2026. The gap is real, growing, and most companies aren’t measuring it at the customer level — where it surfaces before aggregate reporting ever catches it, as a cost-to-serve problem concentrated in a small percentage of accounts.
What’s Actually Happening to SaaS Margins
The driver is inference. AI cost of goods sold now represents 40–50% of revenue for AI-native products, with inference alone often accounting for 20% or more. For traditional SaaS companies layering AI onto an existing product, inference costs represent 4–9% of revenue — a significant new COGS line that often appears with no corresponding pricing adjustment.
What most teams notice is that per-token prices fell 60–75% through 2025. What they underestimate is that those savings are largely offset by the increasing complexity of mature AI features. Retrieval pipelines, multi-step reasoning, and agentic workflows consume far more tokens per user interaction than a simple prompt-and-response ever did. Cheaper tokens don’t fix a more expensive feature.
The net result: AI COGS keeps climbing even as the price per token falls.
What AI Inference Costs Actually Look Like
Inference cost is what you pay a model provider every time your product calls their API to generate a response — billed per million tokens, where a token is roughly 0.75 words. Prices vary significantly by model tier and affect per-customer margin directly.
Current pricing from the two dominant providers (as of May 2026):
| Model | Provider | Input (per 1M tokens) | Output (per 1M tokens) |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 |
| GPT-4o | OpenAI | ~$2.50 | ~$10.00 |
| GPT-4o Mini | OpenAI | ~$0.15 | ~$0.60 |
Batch processing reduces costs ~50%; prompt caching reduces cached input cost by ~90%.
Now apply this to a real interaction. A single AI-assisted feature — a document analysis, a customer-facing summary, a multi-turn conversation — might consume 2,000–10,000 tokens. At Claude Sonnet pricing, that’s $0.03–$0.15 per interaction. If a power user triggers that feature 200 times a month, you’re looking at $6–$30 of inference cost per customer per month, before compute, storage, or retrieval on top.
For a customer on a $99/month legacy plan, that math doesn’t work. This is where AI features go from exciting product capabilities to quiet margin liabilities — and why the companies navigating this well treat AI inference as a managed cost line, not a technology footnote.
Three Ways AI Features Erode Customer Margin
Unbounded usage on flat pricing. When AI features are included in a flat subscription with no usage caps, high-consumption customers generate substantially more inference cost than low-consumption ones — with no revenue difference. Every heavy user is a margin liability at scale. The plan that looks profitable on paper tells a different story once usage is mapped to cost.
Model selection misalignment. Most product teams default to the most capable model available. But not every feature requires Opus-level reasoning. Routing a simple summarization task through a premium model when a cheaper model would produce identical results is a direct margin drain — and it’s largely invisible because model costs aren’t tracked per feature or per customer. What we repeatedly see is that engineering optimizes for output quality in isolation, while finance has no visibility into what that optimization costs per account.
No attribution at the customer level. Most finance teams receive a single monthly AI API invoice. Without mapping token spend to specific customers, features, or usage tiers, there’s no way to know which accounts are profitable and which are eroding margin. AI inference costs only become manageable once they can be allocated to specific customers and pricing tiers — a discipline most finance teams haven’t yet applied to their AI API bill. The aggregate looks acceptable. The distribution rarely does.
A Framework for AI Margin Visibility
The starting point is treating inference cost as a first-class COGS line — not a line item buried inside your cloud bill. Once it’s visible, you can manage it.
The metrics that give real per-customer visibility:
| Metric | What It Measures | Why It Matters |
| AI cost per customer / month | Inference + retrieval spend per account | Reveals which customers are margin-negative on AI |
| AI gross margin by tier | Revenue minus AI COGS by pricing tier | Exposes flat-rate plans subsidizing heavy users |
| Token cost per feature | Average spend per product workflow | Identifies features needing model routing or caps |
| Inference Efficiency Ratio | AI revenue ÷ inference cost | Top-line signal for AI unit economics health |
| ARPU uplift from AI | Revenue increase attributable to AI features | Determines if AI pricing is recovering the cost |
The Inference Efficiency Ratio — AI-attributable revenue divided by inference cost — is the clearest single signal for AI unit economics health. A ratio below 1.0x means your AI spend is generating less revenue than it costs. Well-run AI products typically target 3–5x or higher. Track it alongside AI gross margin as a core metric for any product with embedded AI features.
The operational levers that follow from this visibility: model routing (directing simpler tasks to cheaper models), usage-based pricing or consumption caps on AI-heavy features, and prompt optimization to reduce token volume per interaction. Teams applying these consistently reduce AI COGS by 20–40% without degrading user experience.
On pricing, the clearest structural insight from 2025–2026: the most successful AI SaaS companies use one primary metric the customer understands — outcomes, tasks, or seats — and one internal metric to protect margin. The gap between those two is where AI unit economics either hold or don’t.
In practice, this work starts with understanding which customers are actually driving AI spend — and that requires the same cost attribution discipline that applies to infrastructure, support, and vendor costs more broadly. The companies with the clearest AI margin picture are almost always the ones that have already built structured cost allocation across the rest of their P&L.
The Bottom Line
AI features are not inherently margin-negative. But they are cost-variable in a way traditional software isn’t — and most SaaS pricing and financial reporting structures weren’t built to account for that variability.
The companies navigating this well treat AI inference as a managed cost line, track it per customer and per feature, and price in ways that reflect actual consumption. The ones that don’t are discovering the margin problem quarters after it starts — when it’s harder and more disruptive to fix.
Start by pulling your AI API spend for the last 90 days. Do you know which customers and features drove it? If not, that’s where to begin.
Brad Perry is the CEO of Cogs’z, a profitability management platform built for B2B SaaS companies. Brad co-founded DealerSocket, an end-to-end platform in the automotive industry, where he experienced firsthand the margin challenges that Cogs’z is designed to solve. Cogs’z automates customer-level cost allocation so finance, CS, and sales teams share a single, accurate, view of who’s profitable and why. Learn more or request a demo at cogsz.com.
References
- SaaS Capital — Annual SaaS Gross Margin Benchmarks
- Bessemer Venture Partners — State of the Cloud / AI SaaS Unit Economics, 2025–2026