
We wrote a while back that credit systems are challenging, and that post was about mechanics: grants, rollover, burn-down priority, balances, negative balances. All the things that make a credit system work.
It did not answer the question teams actually get stuck on. What should one credit cost?
That question turns out to be four smaller ones, and getting any of them wrong shows up as a margin surprise a quarter later.
A credit is a unit of value, not a unit of cost
Start here, because it decides everything else. A credit is a denomination you invent. It is worth whatever you say it is worth, and its whole purpose is to sit between two numbers that move independently: what your customer pays, and what your provider charges you.
That gap is the point. When a new model ships or a provider changes its prices, your customer's balance still applies. Nothing gets repriced retroactively and nobody migrates plans. That is the single best property of a credit system, and it only holds if you keep the credit's definition stable.
So define it once, explicitly, and write it down where customers can see it:
One credit = 1,000 tokens of our standard model.
Now you have something to price.
Four numbers
The numbers below are made up to show the shape of the calculation. Use your own.
1. Your cost per credit. If a credit buys 1,000 tokens and your blended model cost is $4.00 per million tokens, a credit costs you $0.004. Blended means input and output weighted by your actual traffic, not the cheaper of the two.
2. Your price per credit. Say $0.010. That is a 2.5x markup and a 60% gross margin on consumption alone. For context, the median target gross margin for AI capabilities across 230 B2B software and AI companies in 2026 was about 50%, against 70 to 80% for traditional SaaS. So 60% is healthy.
3. Your platform fee. The fixed part, if you have one. Say $500 a month for access, independent of consumption.
4. Your included allowance. How many credits the platform fee covers. Say 25,000.
Those four numbers produce your revenue and your cost at any volume:
revenue(c) = platform_fee + max(0, c - included) × price_per_credit
cost(c) = c × cost_per_credit
margin(c) = (revenue - cost) / revenueThe curve nobody plots
Plot margin against consumption and it does something counterintuitive.
At zero consumption your margin is 100%. You collected $500 and spent nothing.
Through the allowance your margin falls, because consumption costs you while revenue sits still. At 25,000 credits you have spent $100 against $500 of revenue, so margin is 80%.
Past the allowance it settles toward the margin on a single credit, which is 60%.
Two consequences fall out of that shape, and both matter.
Your best-margin customer is the one who barely uses the product. Which is also your worst renewal. If you are wondering why the accounts with the healthiest margin churn hardest, that is the curve you are looking at. Worth knowing that 30% of outcome-priced companies and 20% of usage-priced companies report customers deliberately self-policing their spend as a top-two pricing problem. Those customers look great on a margin report and are quietly deciding not to expand.
The allowance goes underwater when it costs more than the fee collects. With a $500 fee and a $0.004 cost per credit, you break even on the allowance alone at 125,000 credits. Set the allowance above that and every customer who uses their full allowance loses you money before a single overage credit is billed.
That break-even number is not a curiosity. It is where your hard ceiling belongs.
Where teams get this wrong
One credit buys everything. A cheap completion and a 200,000-token long-context request should not cost the same number of credits. Long context costs superlinearly. If one credit buys both, your heaviest users are being subsidised by everyone else, and your blended margin will hide it until one account gets big enough to notice.
Price along the dimensions that actually differ in cost: model tier, input versus output, cached versus uncached input, context length, priority class. A credit can cost more for expensive work. That is not complexity for its own sake, it is the difference between a rate card and a wish.
The allowance is set by feel. It is usually the number that produces a comfortable-looking price page. Work out the break-even volume first, then choose an allowance below it.
The markup is set once and never revisited. Provider prices moved in both directions during 2026. One provider halved a consumer tier from $200 to $100 a month; another moved enterprise consumption to raw API rates. Your cost per credit is not a constant, so check what a 40% swing in model price does to your margin before it happens rather than after.
Nobody owns the number. Below $5M ARR the founder usually owns pricing. Past $50M ARR it moves to product or finance. In between it frequently belongs to nobody, which is exactly when a credit rate gets set in a spreadsheet and then forgotten.
What to do with the answer
Once you have the four numbers and the break-even volume, the rest is configuration rather than strategy:
- Publish the credit definition and hold it stable. This is the promise that makes credits tolerable for a customer.
- Set the hard ceiling below break-even, per team and per key rather than one number for the account. Entitlements with metered balances are how that gets enforced.
- Alert well before the ceiling, and send it to the person who owns the budget rather than only the person who owns the API key. They are usually different people.
- Report margin per customer, not blended. At a 50% target the blended number is the average of profitable small accounts and unprofitable large ones, and it will look fine right up until it does not. Cost attribution is the mechanism.
A credit system is not hard to build. It is hard to price. The arithmetic above takes an afternoon, and it is considerably cheaper than finding out from an invoice.


