Solutions | token meter

The bill arrives after the money is gone. A meter doesn't.

Operant Token Meter tracks token consumption as it happens and enforces limits while it's still flowing — down to the exact service, application, team, user, agent, and model doing the spending.

Nobody could say no to AI agents.
Then the tokens started adding up.

The business wanted them, developers wanted them, and the market gave no room to be the person slowing things down. So you said yes — and nothing was watching closely enough to catch what happened next.

Agents going wild on the task

Re-reading an entire codebase to find one file. Pulling the maximum context window on every call. Hammering an endpoint hundreds of times against a transient error. Reasonable once; ruinous at machine speed, thousands of times over.

Behavior that drifts

An agent doing more than it was designed to do, in ways its task can't explain.

An agent steered somewhere it shouldn't go

Rare, and the one that keeps security leaders up at night.

See your AI spend before it becomes a bill.

Operant gives teams a live view of token consumption and the controls to stop runaway usage before it turns into unexpected cost.

Gateway Connected Gatekeeper Connected
0 Active sessions 0 Connected servers
0 / 0
Servers
0
Tools
across 0 servers
0
Sessions
active

Active Servers

C
Cloudflare
STREAMABLE-HTTP · via Custom
just now
A
Angular
STDIO · via Custom
just now
B
Browser Use
STREAMABLE-HTTP · via Custom
just now

Activity log

03-17 21:43:51[INFO]gateway listening on :8080
03-17 21:43:51[INFO]gatekeeper handshake ok
03-17 21:43:57[INFO]gateway registered cloudflare — 24 tools
03-17 21:43:57[INFO]gateway registered angular — 12 tools
03-17 21:44:02[INFO]gateway registered browser-use — 22 tools
03-17 21:44:05[INFO]session opened id=7f3a1c
03-17 21:44:11[INFO]tool catalog refreshed — 58 tools
03-17 21:44:18[INFO]health check passed for 3 servers
03-17 21:44:24[INFO]session opened id=b21e90
03-17 21:44:31[INFO]gateway heartbeat ok
03-17 21:44:39[INFO]tool call search_jira duration=142ms
03-17 21:44:47[INFO]gatekeeper policy sync complete

Two things, and they only work together.

Cost doesn't originate at the account level. It originates with a specific agent, running a specific application, for a specific user, in a specific team, against a specific model.

Granularity across the segments that matter

Service, application, team, user, agent, model — so you can isolate the exact source of the spend instead of staring at a blended total.

■ On its own: a nicer dashboard. You watch the number climb and pay the bill anyway.
Real-time controls at that same granularity

Enforce a limit on that exact segment as the tokens are being consumed, not after they're gone.

■ On its own: a kill switch. You save money by taking down workloads the business needs.
SERVICE
APPLICATION
TEAM
USER
AGENT
MODEL
8.42M
TOKENS
LIMIT · 26%
ServiceALL WORKLOADS
ApplicationSUPPORT-COPILOT
TeamENGINEERING
UserJ. OKONKWO
AgentCONTEXT-LOOP
ModelCLAUDE-SONNET
TOKENS IN METERED LIMIT FIRES AT THE AGENT — FLOW CONTINUES PAST IT
THE LIMIT BINDS ONE RING ONLY
the agent is throttled to 26%; every scope around it keeps flowing untouched

Five ways this actually happens.

What they share: a top-level, after-the-fact number would have missed the cause of every one, and left no way to stop the spend without a blunt shutdown.

Coding agents across a few hundred engineers
A handful were configured to re-read large repositories on every step, and those few quietly ran up most of the month's token spend. The account-wide total showed the number climbing but never pointed to which agents drove it.
A real-time per-agent cap would have throttled those specific agents on day one and cut the bill by the majority, while the other few hundred engineers kept coding, untouched.
A retrieval pipeline with no bound on context
On certain queries it expanded context windows to the maximum on every call, looping as it went. Each request looked reasonable on its own; in aggregate it spent several times what the workload warranted.
A real-time per-application limit would have capped that one pipeline before the overage accumulated, without touching the rest of the data stack.
An evaluation suite wired to a live model endpoint
It ran on every commit and worked perfectly — burning production tokens continuously for two weeks before anyone connected the invoice to the CI job.
A per-service token limit would have flagged and throttled that endpoint within hours of the spend leaving its normal band, turning a two-week leak into a same-day catch.
A customer-facing agent falling back to the largest model
During a traffic surge it sent every request, including the simple ones, to the most expensive model. With no per-model policy, cost per interaction jumped an order of magnitude for days.
A real-time per-model cap would have kept the expensive model reserved for the requests that actually needed it.
An agent whose consumption spikes in a shape its task can't explain
A bug, or occasionally an agent steered somewhere it shouldn't go. The rarest case, and the one security leaders ask about first.
Real-time metering turns that into a limit that already fired, on that one agent, before it could run up either the bill or the risk.

Three capabilities, one surface.

Operant Endpoint Protector is purpose-built for the threat model of coding agents — not adapted DLP, not bolted-on API gateway, but designed from the ground up to operate at the only layer where coding agent risk can actually be prevented: the developer's device.

See usage as it happens
Near real-time token metrics per agent session, broken out by user, team, agent, and model. Usage insights, recommendations, and block metrics.
Enforce budgets at runtime
Token limit policies that constrain user and agent activity even mid-session. Granular limits by user, team, agent, and model, plus organization-wide caps.
Cover every deployment
All models supported, including the hosted options through Bedrock, Vertex, and Foundry that providers leave dark. One policy spans every agent and flavor.

Meter every token.
Control every workload.

You shouldn't have to choose between giving engineers powerful AI tooling and keeping the bill under control. The million-dollar bill isn't the price of using AI. It's the price of not being able to meter it.