Latest
Daily Tech Times Subscribe

Unit economics for AI products: pricing per token, query or seat against real compute cost

Before a founder can talk about gross margin as a percentage, they need to know what it actually costs to serve one customer doing one thing, and UK AI startups are still working out the best way to measure that.

black and yellow rubber puzzle mat
Photo · Photo by Ryan on Unsplash

Why the unit matters more than the percentage

Most startup metrics conversations jump straight to gross margin as a headline number. For an AI company that is putting the cart before the horse. Gross margin is just the output of a calculation, and the calculation depends entirely on what unit you choose to measure. A SaaS business built on a database has a fairly stable cost per seat. An AI product built on a large language model has a cost that moves with every prompt, every token generated, every image rendered or every minute of audio transcribed. Founders who skip straight to a margin percentage without first nailing the unit economics underneath it are often working from numbers that will not survive contact with real usage.

The starting point is choosing the right unit: per API call, per token (input and output separately, since they are usually priced differently by model providers), per active seat per month, or per completed task. Different products suit different units. A coding assistant might think in tokens and completions. A customer support tool might think in resolved tickets. A voice product might think in minutes processed. Getting this wrong means a founder can report healthy blended margins that hide loss-making heavy users, the ones most likely to expand their usage and eventually blow up the model.

Building the real cost stack

Once the unit is chosen, the next job is listing every cost that scales with it, not just the obvious one. The headline cost is usually inference: calling a foundation model API, or running a self-hosted model on rented GPU capacity. But a full stack usually includes retrieval and vector database costs if the product uses retrieval-augmented generation, embedding generation, orchestration and logging infrastructure, content moderation or safety checks, storage of conversation history, and any fine-tuning or evaluation runs that get amortised across usage. Founders who only cost the headline model call typically understate true cost per unit by a meaningful margin, sometimes enough to turn a product that looks profitable at the unit level into one that is not.

It also matters whether a startup is calling a third-party model API or running its own infrastructure. API-based products have highly variable, transparent-ish costs that move when the provider changes pricing, which happens more often than founders would like. Self-hosted or fine-tuned open-weight models can have lower marginal cost per call once utilisation is high, but carry fixed costs in GPU reservation, DevOps and model maintenance that only make sense at volume. Many startups start on APIs to move fast and validate demand, then consider bringing parts of the stack in-house once volume is proven and predictable.

Why blended margin numbers mislead investors and founders alike

A single blended gross margin figure for an AI company can hide enormous variation between customer segments. A light user who asks a handful of questions a day might be hugely profitable. A power user who runs the product continuously, or who feeds it very long documents, might cost more to serve than they pay. Because AI usage patterns are so uneven, founders need cohort-level unit economics, not just a company-wide average, before they can credibly claim a margin trajectory. Investors doing diligence increasingly ask for this breakdown specifically because blended numbers have been used to paper over concentration risk in a handful of expensive users.

This is also why pricing model choice is inseparable from unit economics. Flat per-seat pricing is simple to sell but dangerous if usage is not capped, because the heaviest users generate the most cost against a fixed price. Usage-based pricing (per token, per call, per credit) aligns revenue with cost more closely but can be harder for customers to budget for and can slow sales cycles. Many AI startups end up with a hybrid: a base seat fee that covers a usage allowance, with metered charges above it, precisely because it lets them protect margin on heavy users while keeping the product simple for light ones.

What good practice looks like

A founder who understands their unit economics well can usually answer a few specific questions without hesitation: what does it cost to serve the median customer for a month, what does it cost to serve the 95th percentile customer, what proportion of revenue at the current price point is consumed by underlying compute for each segment, and how does that ratio change as model providers change pricing or as the company migrates workloads to cheaper infrastructure. None of these numbers are fixed industry constants, model API pricing, GPU rental rates and cloud provider discount structures all change over time, so founders should track them against current supplier pricing rather than relying on figures from when the product first launched. Getting the unit economics right early makes every conversation that follows, about pricing, about fundraising, about margin trajectory, considerably more grounded.

Sources