
Outcome-based pricing sits at 5% of B2B software companies today. Asked where they expect to be in three years, respondents in the 2026 Growth Unhinged survey put it at 31%.
The pitch behind that jump is genuinely strong. You only pay when it works. It removes the "does this actually deliver" objection in one sentence, which is why several large vendors shipped a version of it in the first half of 2026.
We compared usage-based and outcome-based pricing last year and covered how the models relate. This post is about the side of outcome pricing that does not appear on the pricing page: what the failures cost you.
The arithmetic
The resolution rates and per-attempt costs below are made up to show the mechanism. Substitute your own measured figures before making a pricing decision on them.
Say you charge $0.50 per resolved ticket and an attempt costs you $0.12 in inference. At a 70% resolution rate:
- 100 attempts cost $12.00
- 70 resolve, earning $35.00
- Gross margin: 66%
Comfortable. Now a customer onboards with a messier knowledge base than your design partner had, and the resolution rate drops to 45%:
- 100 attempts still cost $12.00
- 45 resolve, earning $22.50
- Gross margin: 47%
Still workable. Now add the part that actually happens. Failed attempts are not cheap attempts. An attempt that fails usually fails after the agent has tried several approaches, expanded its context and called more tools. Say a failure costs 2.5x a success:
- 45 successes at $0.12: $5.40
- 55 failures at $0.30: $16.50
- Total cost $21.90 against $22.50 of revenue
- Gross margin: 2.7%
Same published price. Same product. The entire difference lives in two numbers that sit in your agent's control loop rather than in your pricing model: how often it fails, and what a failure costs relative to a success.
If you do not measure both, you do not know your outcome-based margin. You know your outcome-based revenue, which is a more flattering number.
Four things to build before you publish the price
1. A step budget per task
The most common cost runaway in agentic products is a loop that retries until it succeeds, because "keep trying" looks like a sensible default. Cap iterations per task, tool calls per iteration, and total tokens per task.
When a task hits the ceiling it should fail deliberately and cheaply, rather than failing expensively an hour later. A deliberate failure at a known cost is something you can price. An open-ended attempt is not.
2. A model-routing policy you can see
Cheap tasks must not quietly execute on expensive models. Route by task class and alert when the routing drifts, because it will, usually after someone changes a default to fix a quality complaint and nobody tells finance.
This is also where cost attribution earns its keep. If you cannot see cost per customer per model, a routing change that halves your margin looks like a good week until the invoice arrives.
3. A failure taxonomy, written down before the first dispute
"Failed" is not one thing, and billing as though it were is how you end up arguing with a customer during a renewal. At minimum separate:
| Case | Who pays | Why it matters |
|---|---|---|
| Agent could not do it | You. No charge. | This is the case the pitch is about |
| Bad or missing input | Arguable | Measure it separately, it is coachable behaviour |
| Upstream provider failed | You, and retry for free | Not the customer's problem |
| Customer cancelled mid-task | Decide now, in writing | You did the work |
| Resolved, then reopened | The hard one | Is that a failed resolution or a new ticket? |
Pick the definitions before your first enterprise renewal. The customer will have opinions and they will have them at the least convenient moment.
4. Per-attempt enforcement, not per-invoice reporting
This is the structural one, and it is why the other three are hard to run from downstream tooling.
Your billing system sees resolved outcomes, because resolved outcomes are what you bill for. It does not see the 55 failures, because those never became line items. So the component that knows your revenue cannot see the majority of your cost.
The costs are incurred at the model call, which means the controls have to live there too: step caps, routing policy, and an alert to you rather than the customer when a single account crosses its margin floor. That is the control that catches a negative-margin account in week one instead of at quarter close.
The four-condition test
Kyle Poyar's CAMP framework is a useful gate before committing to outcome pricing. All four conditions have to hold:
| Condition | The question | You fail it when |
|---|---|---|
| Consistency | Are outcomes reliable and repeatable? | Resolution is 40% one week and 70% the next |
| Attribution | Is the result clearly yours? | A human touched it, or three tools share credit |
| Measurability | Can it be quantified concretely? | "Improved satisfaction." Nobody can audit that |
| Predictability | Can the customer forecast outcomes and cost? | They cannot put a number in next year's budget |
Predictability is the one that loses deals, and it is the one vendors forget because it is the customer's problem rather than theirs. You can have consistent, attributable, measurable outcomes and still lose to a flat-rate competitor because procurement cannot forecast the bill.
The fix is not abandoning the model. It is wrapping the outcome metric in a committed annual amount with a floor and a ceiling. The customer gets a number they can plan against and you keep the "only pay when it works" story. That is also why 29% of companies now let customers choose between pricing models rather than betting the business on one.
The short version
Outcome pricing sells your successes and buys your failures. Before you publish the price, know what a failure costs, how often it happens, who caused it, and where the ceiling sits.
Otherwise you have written an appealing sentence on a pricing page and handed your gross margin to a retry loop.


