
An agent gets stuck in a retry loop overnight. It is a reasonable loop, because the task keeps almost succeeding, so it keeps trying. By morning it has spent a month of token budget for an organization of four hundred people.
Your billing system has every event. Correctly aggregated, correctly priced, ready to produce an accurate invoice. What it could not do is stop it.
That is not a missing feature. It is a property of where the component sits.
Downstream systems are built to be correct about the past
A usage-based billing platform sits downstream of your application by design. Your code emits events, the platform ingests them, aggregates, applies a rate card and produces an invoice.
That architecture is right for accounting, and accounting is a job about the past. Ask the same component to make a decision about the present and you have asked the wrong part of the system. By the time the event lands, the tokens are spent and your model provider has already billed you.
So a "spending limit" configured downstream is worth looking at closely. What you configured was a notification with a threshold attached. Useful, but not a limit.
The two questions are asked half a second apart
There is a second cost to the downstream shape, and it shows up in how this tooling is usually assembled. Metering and entitlements tend to arrive as separate concerns, often separate vendors, answering two questions:
- How much did they use?
- Are they allowed to use it?
Those are the same question with a different tense. Splitting them across two systems means two integrations, two balances and two opportunities to disagree. Under normal traffic nobody notices. Under a retry loop firing hundreds of concurrent requests, they all read the same stale answer.
This is why we built entitlements into the same system that does the metering rather than beside it. Not because it is tidier, but because the balance you check and the balance you increment should not be two different numbers maintained by two different services.
What changes when the decision is in the path
Every API call and every model call already passes through something that authenticates it and routes it. That component knows the organization, the team and the key. Which means the thing that counts the usage can also refuse it.
Four consequences:
A limit becomes a limit. The check runs before the request is served. At zero balance the next call is refused, before the token is spent.
One balance instead of two. The check and the increment belong to the same system, so there is no second source of truth to fall behind.
It covers code you do not own. Model providers, MCP servers, partner APIs, an agent calling five tools you did not write. You cannot instrument somebody else's service. You can proxy it.
The decision lands where the cost is incurred. Token cost happens at the model call. That is the only place in the system where stopping spend before it becomes spend is available at all.
The part we should be precise about
Enforcement in the request path has a latency cost and a failure mode, and anyone who tells you otherwise is selling something.
Our Edge Access work exists for exactly this. Balances and entitlement values live in a globally distributed dataset, updated on every usage event and every entitlement change, with reads routed to the nearest region at around 25ms.
Fast is not the same as current. The metered balance can take several seconds to reflect the most recent events, so a burst of concurrent requests inside that window can read a number that has not caught up.
And the caveat from that post is worth repeating plainly: those reads are eventually consistent. The metered balance can take several seconds to reflect the most recent events. A burst of concurrent requests inside that window can read a number that has not caught up yet.
So the claim is not that a consistency window disappears. The claim is narrower and still worth a lot:
- The window goes from a billing cycle to a few seconds.
- The decision exists, which it did not before.
- The request can be refused rather than reconciled.
Whether you fail open or fail closed inside that window is a policy decision per feature. Blocking is the right default for an expensive long-context call and the wrong default for a health check. Pick deliberately, and know what your system does when it cannot reach a balance at all.
One runaway user should not drain the account
Most spend controls are a single number for the whole account, which makes every user a shared-fate risk. That is how one overnight loop ruins four hundred people's morning.
The version that helps puts limits at each identity you can already see on the request. In OpenMeter the subject is a key carried on the usage event, so a metered entitlement can sit at the organization, the customer, the team or the individual key.
Then the loop hits its own ceiling, the agent stops, and everybody else keeps working.
Why this is a growth feature and not just a safety feature
The survey data on this surprised us. Among companies pricing on outcomes, 30% named "customers self-police usage and spend" as one of their top two pricing problems. For usage-priced companies it was 20%.
Their customers are deliberately using less than they want to, because they do not trust what the bill will look like.
Limits get sold as margin protection, and they are. But a limit the customer can see, set and rely on is also what lets them stop rationing. If you are trying to grow consumption, the fastest thing you can do is make the ceiling visible.
Start with entitlements and balances and grants.


