
Trace a single agent run in a production system. A user asks for something, and the agent calls a model provider, queries two MCP servers for context, hits a partner API for enrichment, calls the model again to summarise, and writes a result.
Six billable interactions. You wrote one of them.
Every metering approach that starts with "add our SDK to your service" quietly assumes you own the service. In an agentic architecture that assumption has stopped holding, and it is the reason agent costs are so hard to attribute.
Instrumentation does not reach other people's code
The SDK model works when your application is the thing doing the spending. You import a client, emit an event per operation, and the meter adds it up.
Now list what an agent run spends money on:
| What the agent calls | Who wrote it | Can you add an SDK? |
|---|---|---|
| Model provider API | Not you | No |
| MCP server, third party | Not you | No |
| MCP server, internal | You | Yes, and now you have two systems |
| Partner or vendor API | Not you | No |
| Your own service | You | Yes |
You can instrument the last two rows. Those are usually the cheapest rows.
The expensive spending happens in code you cannot deploy to, which leaves you reconstructing cost afterwards from provider invoices that arrive monthly, aggregate by API key rather than by customer, and cannot tell you which agent run caused what.
What the proxy hop sees
There is one place every one of those six calls already passes through, if you put it there: the gateway in front of them.
A proxy hop does not need cooperation from the thing on the other end. It sees the request and the response, which is everything metering needs:
- Identity on the way in. Which customer, which team, which key, which session. This is the part provider invoices can never give you.
- Route and upstream. Which provider, which model, which MCP server, which tool.
- The payload. Token counts from the response body, or from usage fields the provider returns. We covered the mechanics for OpenAI-style APIs and for streamed responses, where the usage block arrives in the final chunk.
- Outcome and latency. Whether the call succeeded, which matters a lot if you price on outcomes.
That is a usage event with customer attribution attached, produced without touching the upstream service.
The correlation problem, which is the actual hard part
Metering each hop is the easy half. The half that decides whether your numbers are useful is tying six hops back to one agent run and one customer.
Two things make that work, and both belong at the proxy:
A run identifier that survives every hop. Propagate a correlation header from the first request through every downstream call the agent makes. Without it you have six unrelated usage events. With it you have one agent run that cost $0.41, and you can answer why.
Identity resolution at the edge. The gateway knows the authenticated principal. Stamp the customer and team onto every derived call, so a tool call made three hops deep is still attributable. In OpenMeter terms the subject is a key on the usage event, so whatever the gateway resolves is what you can meter and, later, enforce on.
Get both and cost attribution stops being a monthly reconciliation exercise. You can already attribute shared cost to customers and internal teams; this is the same idea applied to spending that leaves your perimeter.
MCP is a billable surface nobody is counting
Model Context Protocol servers are becoming a standard way to give agents tools, and a tool call is a billable event in both directions.
If you operate MCP servers, each call is consumption someone should be accountable for, and some tools are far more expensive than others. A cheap metadata lookup and a query that scans a warehouse table should not cost the same number of credits.
If you consume third-party MCP servers, each call is cost you are absorbing on a customer's behalf, and you currently have no per-customer view of it.
Either way the meter goes at the hop, because there is no other shared point. And once it is there, the interesting capability is not the report. It is that a tool call can be refused.
A budget that spans services you do not own
This is the part that only works in the path.
A per-run budget means the agent gets a ceiling for the whole task rather than one limit per service. The run starts with a budget, each hop draws it down, and when it is exhausted the next call is refused, whether that call goes to your service, a model provider or somebody's MCP server.
You cannot build that from provider-side limits. Each provider enforces its own quota in its own units and none of them knows about the others. The only place with a view of the whole run is the hop they share.
Practically, three controls do most of the work:
- A run budget in credits, drawn down across every hop.
- A step cap, because the failure mode is a loop rather than a single expensive call.
- Per-tool pricing, so an expensive tool costs what it costs. This is what entitlements and metered balances are for.
The short version
Metering used to be something you added to your code. In an agentic system, most of the spending is not in your code, so metering has to move to the hop that every call shares.
That hop also happens to be the only place that can stop a call before it spends anything. Which makes it the natural home for both halves of the problem.


