Why AI startup gross margins start low and what a credible path back up looks like
AI companies often post gross margins well below the SaaS benchmarks investors are used to, so founders need a clear story for why that changes over time.
Why this matters beyond the headline number
Most UK investors were trained on SaaS economics: a well-run software business should eventually post gross margins in the high 70s or 80s percent, because once the code is written, serving one more customer costs almost nothing. AI products break that assumption. Every user query, every generated image, every agent call runs through a model that consumes compute at the moment of use. That is a real, variable cost that scales with usage in a way traditional software costs did not.
This means a growing AI startup can look, on paper, like it is getting worse at unit economics even as it adds customers, simply because inference cost is a genuine cost of goods sold, not a one-off engineering cost. Founders who understand this, and who can explain why their margin improves over time, raise more easily than founders who just hope investors do not ask.
The three levers that actually move margin
Gross margin in an AI business is mostly a function of three decisions, not of growth rate alone.
The first is model choice. Calling a frontier model through an API for every request is simple to build but expensive to run, because the startup pays someone else’s margin on top of the raw compute. Fine-tuning a smaller open-weight model, or routing simple queries to a cheaper model and only escalating complex ones to a larger one, cuts cost per query substantially. This is why so many teams now build routing layers rather than hard-wiring themselves to a single model provider.
The second is infrastructure engineering: caching repeated answers, batching requests, quantising models, and negotiating committed-use discounts with cloud or model providers. None of this is glamorous, but it is often where the actual margin improvement comes from between one funding round and the next, more than any change in the product itself.
The third is pricing architecture. Usage-based pricing passes compute cost variability straight to the customer, which protects margin but can make revenue lumpy and harder to forecast. Flat subscription pricing is easier to sell and forecast but leaves the startup exposed if usage patterns shift, particularly with power users who consume far more compute than an average customer. Many AI companies now blend the two: a subscription with usage caps, or seat-based pricing with metered overage, precisely to manage this risk.
What investors are actually checking for
At seed stage, investors generally accept low or even negative gross margin as long as the founder can articulate why it improves. What they are checking is whether the cost structure is understood at all: does the founder know the cost per query or per customer, and can they show it trending down as volume, caching, or model efficiency improves. A founder who cannot answer “what does it cost us to serve our average customer for a month” is a red flag regardless of growth rate.
By Series A and beyond, the bar rises. Investors want to see a credible glide path towards SaaS-like margins, even if the business never quite gets there, because it changes how much capital the company needs to reach profitability and how it should be valued. A business with 40% gross margin needs a fundamentally different amount of cash to scale than one heading towards 75%, and that difference compounds at every later round.
It is also worth separating two things that get conflated: the cost of running the product for existing customers, and the cost of training or fine-tuning models in the first place. Training cost is more like R&D or capital expenditure, it is spent once, or periodically, to build or improve the model. Inference cost is the recurring cost of serving each customer request. Blending the two makes gross margin look worse or better than it really is, and sophisticated investors will ask founders to separate them cleanly.
Practical steps founders can take
Track cost per query or cost per active customer as a core metric from an early stage, not as an afterthought before a fundraise. Model out how that cost changes as volume scales, distinguishing improvements that come automatically from scale (better unit pricing on compute) from those that require engineering work (caching, model distillation, routing). Build pricing so that the most compute-hungry customers pay proportionately more, rather than subsidising them out of margin earned from lighter users. And when talking to investors, present gross margin as a trajectory with named drivers, not a static snapshot, because that is the framing that actually answers the question they are asking.
For the current detail on R&D tax treatment of compute and cloud costs, and on any government-backed compute schemes relevant to cost planning, check the primary sources directly rather than relying on a fixed figure, since these terms are reviewed periodically.