

There's a bill that a lot of engineering and security leaders are dreading right now. It shows up after the month closes, it's far bigger than anyone forecast, and no one can fully explain it. It's the AI bill (cue scary dramatic music, dun dun dun) — and the reason it landed on your desk is that, quite reasonably, you couldn't say no to AI agents. The business wanted them, developers wanted them, and the market gave you no room to be the person slowing things down. So you said yes. And then the tokens started adding up in ways nobody was watching closely enough to catch.
Today we're introducing Operant Token Meter to close that gap. Operant now provides deep token tracking and controls: metrics and insights alongside policies with real rate-limiting and model usage controls for Claude and Codex deployments across your organization. It meters token usage in real-time — reining in costs and rogue agents as they happen, at the exact level they happen — without reining in the rest of your AI.
The uncomfortable truth about that seven-figure surprise is that it's rarely one dramatic event. It's the accumulation of agents quietly doing far more work than their tasks required, for weeks, with nothing in place to stop them.
By far the most common cause is simple: an agent goes wild solving the problem it was given. Not maliciously, not through some exotic emergent breakout — just inefficiently, at machine speed, thousands of times over. An agent re-reads an entire codebase on every step when it needs a single file. A retrieval chain pulls the maximum context window on every call because nobody capped it. A retry loop hammers a model endpoint hundreds of times against a transient error. Each individual action looks reasonable. Multiplied across a fleet of agents running around the clock, it's how a modest pilot turns into a million-dollar line item.
There are other causes too, and they matter — an agent whose behavior drifts past what it was designed to do, or, rarely, one that's been steered by malicious intent. Those are real and this product covers them. But in August 2026, the thing actually running up your bill, day in and day out, is agents burning tokens on expensive, unnecessary ways of getting their job done. That's the fire. And a top-level monthly report is a smoke alarm that only goes off after the house has burned down.
Model providers do offer some visibility, but it comes with real limits. Deeper usage reporting often sits behind higher pricing tiers. What you do get isn't real-time. Coverage is first-party only, so there are no metrics for hosted deployments running through AWS Bedrock, Google Vertex, or Azure Foundry. And there's no enforcement or limits when usage runs past budget — you find out after the fact, once the charges have already incurred.
At their best, the providers' tools stop at telling you a number. That's a bill, not a meter, and a bill can't stop a cost overrun — it can only confirm one.
Agent cost doesn't originate at the account level. It originates at the level of a specific agent, running a specific application, on behalf of a specific user, inside a specific team, against a specific model. That's where the tokens get spent, and it's the only level at which you can do anything about it. An account-wide monthly figure can't tell one agent from a thousand others, so the only lever it offers is the crude one: let everything run, or shut everything off. No leader wants that choice — which is exactly why the spending keeps running.
Controlling agent cost takes two things, and they only work together:
Granularity across the segments that matter — service, application, team, user, agent, and model — so you can isolate the exact source of the spend instead of staring at a blended total.
Real-time controls at that same granularity — the ability to enforce a limit on that exact segment as the tokens are being consumed, not after they're gone.
Granularity without enforcement is just a nicer dashboard — you watch the number climb and pay the bill anyway. Enforcement without granularity is just a kill switch — you save money by taking down workloads the business needs. Neither one controls cost on its own. What controls cost is the two together: fine-grained visibility and fine-grained, real-time enforcement across the same segments. That's what a top-level provider report structurally cannot do, and it's what Operant Token Meter is.
There are some common scenarios that keep coming up as the industry leaps forward with agentic deployments doing more and more complex, business-critical workflows. What they share is that a top-level, after-the-fact number would have missed the cause of every one — and left no way to stop the spend without a blunt shutdown.
A platform team rolled out an agentic coding tool to a few hundred engineers. A handful of agents were configured to re-read large repositories on every step, and those few quietly ran up most of the month's token spend. An account-wide total showed the number climbing but never pointed to which agents drove it. A real-time per-agent cap would have throttled those specific agents on day one — and cut the bill by the majority — while the other few hundred engineers kept coding, untouched.
A retrieval pipeline, on certain queries, expanded its context windows to the maximum on every call, looping as it went. Each request looked reasonable on its own; in aggregate it was spending several times what the workload warranted, because nothing bounded the context growth. A monthly report would have graphed the total going up. A real-time per-application limit would have capped that one pipeline before the overage ever accumulated, without touching the rest of the data stack.
A team wired an agent's evaluation suite to a live model endpoint and ran it on every commit. It worked perfectly — and it burned production tokens continuously for two weeks before anyone connected the invoice to the CI job. A per-service token limit would have flagged and throttled that endpoint within hours of the spend leaving its normal band, turning a two-week leak into a same-day catch.
A customer-facing agent handled a surge of traffic by falling back to the largest, most expensive model for every request, including the simple ones. Nobody had set a per-model policy, so the fallback ran unchecked and the cost per interaction jumped an order of magnitude for days. A real-time per-model cap would have kept the expensive model reserved for the requests that actually needed it.
And the case that keeps security leaders up at night even though it's the rarest: an agent whose token consumption suddenly spikes in a shape its task can't explain — a bug, or occasionally an agent steered somewhere it shouldn't go. Real-time metering turns that into a limit that already fired, on that one agent, before it could run up either the bill or the risk.
Being told you owe a large sum at the end of the month isn't a measurement you can use (other than as an excuse for why you’re hiding under your desk until your boss leaves the room). By the time it lands, the spend is gone, whatever consumed it already ran, and all that's left is to argue internally and try to reconstruct what happened three weeks ago.
Metering is different. It means seeing consumption as it happens, attributing it to the exact service, application, team, user, agent, and model responsible, and enforcing against it in real time — reining in the cost or the rogue agent while it still matters, and while everything that isn't the problem keeps working.
See usage as it happens. Near real-time token metrics per agent session, broken out by user, team, agent, and model. Usage insights, recommendations, and block metrics let you operationalize Token Ops at scale. This is the granular, real-time visibility a top-level provider report can't give you, and it's the foundation the controls sit on.
Enforce budgets at runtime. This is the meter. Token limit policies that constrain user and agent activity — even mid-session — when budgets are exceeded. Granular limits by individual user, team, agent, and model, plus organization-wide caps. Not an alert after the fact: real-time enforcement that acts on the token flow while it's still flowing, and only on the segment that's over the line — the agent or workload running up the cost, not the ones around it.
Cover every deployment. All models supported, including the hosted options through Bedrock, Vertex, and Foundry that providers leave dark. A single policy spans every agent and flavor, from endpoint to cloud, so you're metering one surface at full granularity instead of stitching together partial, coarse views.
Crude tooling that does try to enforce gives you one blunt lever: hit an account-wide number, kill everything. That's more like a breaker switch than a real control, and it takes down your healthy workloads along with the one that went wrong. In practice those limits get set impossibly high or ignored, so they protect nothing — and the security leader is right back to the choice between unchecked usage and a full shutdown.
The answer isn't a better breaker. It's controls that live at the same granularity as the usage. Operant meters and enforces at every level that maps to how your organization actually runs: the service, the application, the team, the individual user, the specific agent, the model. That granularity is what lets you rein in the exact problem without reining in the business. When one agent's consumption runs away, you throttle that agent — not the fleet. When one team's experiment gets expensive, you cap that team — while everyone else keeps shipping. When a single agent goes rogue, you cut it off — in real-time — and everything else keeps moving. The problem gets reined in; the rest of your AI keeps working.
At Operant, we love analogies, and looking out at my garden in the peak of August (albeit, breezy San Francisco August), the use of the term “meter” here has a pretty clear parallel to how good versus bad water metering works.
Picture two houses, both with a "water meter."
The first only measures, and only at the property line. At the end of the month it tells the homeowner they owe $2,000 for the water the new landscaping sucked up. Accurate, unarguable, and useless for changing the outcome — the yard is already overwatered, the money already spent, and the number can't even say which part of the irrigation piping had a leak. That "meter" is really just a bill with extra steps. That's a top-level token report: one big figure for water you've already lost.
The second one is a smart meter that actually meters. It sits inline on the flow, broken out zone by zone, and regulates each one in real-time. The homeowner sees the single sprinkler line that's been running all night — or a pipe that's buried in sandy SF soil while oozing huge amounts where it can’t be noticed by the naked eye — and throttles that one zone. The rest of the garden keeps getting watered. Same landscaping, far smaller bill, and nothing shut off that shouldn't be.
That second “smart” meter is the Operant Token Meter. It doesn't hand you a total after the money's gone, and it doesn't kill the whole system when one number crosses a line. It meters the flow zone by zone — each service, application, team, user, agent, and model — reining in the one running hot, whether that's a cost leak or a rogue agent, while everything else keeps running. It waters the plants instead of cutting the main.
The point of all this isn't to slow teams down or make AI feel like something to ration nervously. It's the opposite: to let adoption keep accelerating because the cost is finally metered and enforced live — precisely enough to stop the overspend without stopping the work. You get to keep saying yes to AI agents, because saying yes no longer means signing a blank check.
You shouldn't have to choose between giving your engineers powerful AI tooling and keeping the bill under control. Token Meter is how you get both: usage metered in real time, across the segments that matter, attributed to the exact agent and team that caused it, and enforced the moment it crosses the line — as the tokens are actively being used, without reining in your AI.
The million-dollar bill on your desk isn't the price of using AI. It's the price of not being able to meter it. Let's put a real meter on it.