Repository navigation
Per-model rate limiting and cost budgets #390
Description
Activity
Shall I work on this features
you should control your token spend limits directly via the provider
all the providers I used have a limit settingthe token spend calculations a harness does might not reflect exactly the provider costs as prices change, extra costs could be added, cache misses, tool calls, etc..
Correction — 24 September 2026. I am the author/maintainer of AEGIS.
I withdraw the claims about physically killing threads, <0.005 ms full processing, and replacing Zero's shared or distributed rate/budget controls. AEGIS is Python; the later rpm-only Go experiment discussed in this issue is separate and is not evidence of AEGIS adoption.
AEGIS Core 3.4.0 provides a Python process-local authorization gate. The caller must supply a known amount, reuse the logical call ID and execute the tool only when execution_permitted is true. A cached authorization receipt is not the tool's result. This does not establish multiprocess coordination, crash recovery or exactly-once external effects. Current scope: https://github.com/LortuArte/aegis-sdk.
Short answer: no, Zero cannot do any of this today. There is no rpm limit, no tpm limit, and no cost cap anywhere in the tree. I grepped for the obvious spellings (
rate.Limiter,x/time/rate,RequestsPerMinute,TokensPerMinute,CostLimit,MaxCost,SpendLimit,budgetUSD) and got zero hits. Not a config key, not a flag, not an env var.Sorry this sat so long. Here is what is actually there, because a good chunk of the work is already done and it is not the part people expect.
The accounting layer is real and finished
This is the good news.
zeroruntime.Usage(internal/zeroruntime/types.go:158) is a normalized per-request record that all three first-party providers emit, with a documented invariant worth reading before you touch it: cache-read and cache-write are subsets ofInputTokens, not additive, andReasoningTokensis a subset ofOutputTokens.modelregistry.CalculateCost(internal/modelregistry/cost.go:38) already turns one of those into a USD breakdown including tiered pricing and the cache splits.usage.Tracker(internal/usage/tracker.go) already aggregates it per model, andSummary().ByModelalready answers "N tokens and $X on model M".So nobody needs to build counting or pricing. One catch: the tracker is constructed only at
internal/tui/model.go:912andinternal/tui/btw.go:112. It is TUI-only and in-memory. Headlesszero execnever builds one at all.Exactly one thing enforces a budget today
Goal.TokenBudget, viaStore.AddGoalUsage(internal/sessions/goal.go:280). It accumulates tokens, compares against the budget, flips the goal toGoalStatusBudgetLimitedwith reason "token budget reached", and emits an event. That is the shipped shape of "a budget that stops work" in this codebase, and a new budget should mirror it rather than inventing a second state machine.But look at where it is called from. One non-test caller:
internal/tui/goal.go:325, insidereconcileGoalAfterRun, which runs after a whole run has finished and sums the run's usage events. So it cannot interrupt anything. A single run overshoots the budget by however much it wants, and the budget only gates whether the next autonomous continuation starts. It is also TUI-only, so exec, ACP and cron never account it at all.Which is a concrete instance of exactly the failure the thread already flagged.
The enforcement point
The warning upthread is correct, and I would put it more strongly: there is nothing to make a network call to. Any limit Zero enforces is a local counter, so the question is not "which service do we ask" but "which local seam sees every request before it goes out".
That seam exists and is purpose-built.
zeroruntime.TurnSession(internal/zeroruntime/session.go:47) is opened once per run, and the loop reassigns its provider local so everything flows through it. Frominternal/agent/loop.goat the assignment:Route ALL provider I/O, per-turn streams, mid-stream retries, the compaction summarizer, and the final-answer request, through the session by reassigning the loop's provider local. Every existing call site reads this local, so no signatures change anywhere downstream.
That is the place. A limiting
TurnSessionProviderwrappingzeroruntime.NewProviderTurnSessionProviderneeds no signature change anywhere, andproviders.OptimizedTurnSessions(internal/providers/turn_session.go:50) is the existing example of installing one behind a gate. Blocking insideStreamis in-process and pre-send, so a runaway loop stops on request N, not request N plus whatever the round trip let through.The thing not to use is
agent.Options.OnUsage(internal/agent/types.go:417). It fires for every request including the compaction summarizer, which makes it the right place to account, but it isfunc(Usage)with no return and no error. It cannot refuse anything. Wire accounting there, enforcement in the session.There is a lower option,
providerio.SendWithAuthRetry(internal/providers/providerio/auth.go:31), where theSpanProviderQueueblock already carries this comment: "No send-side semaphore exists today; a future request queue would also accumulate here." That is provider-agnostic and gets tracing for free, but it does not know token counts, so it only really serves rpm.What Zero can and cannot honestly know
This is where I would push back on parts of the ask, because the three limits are not equally knowable.
rpm is exact. A request count is known locally and precisely before the send. This is the one that works properly, and it is also the one that actually stops a runaway loop.
tpm is approximate going in and unknown coming out. The only token estimator in the repo is
estimateTokens(internal/agent/compaction.go:124), and its own doc says it "deliberately uses no real tokenizer; it only needs to be monotonic and roughly proportional". It is characters over four. Output tokens are not known until the response arrives. So a tpm limit can only honestly be "refuse the next request because the trailing window is already over", never "hold the window under N".cost is a reconstructed estimate, and sometimes it is fiction. Rates are hardcoded in
internal/modelregistry/catalog.gowithsourceLastVerified = "2026-06-04", optionally overlaid from models.dev. Unpriced or unknown models contribute tokens and no cost, by design. And for anyone using a ChatGPT or Claude subscription through the local-proxy recipe indocs/oauth-subscriptions.md, a per-token dollar figure does not correspond to anything they are billed. A cost cap has to say out loud that it is an estimate, the wayzero usage reportalready does.Smallest design that is honest
Config: a new top-level
limitskey onFileConfig(internal/config/types.go:352), keyed by resolved model id. Not onProviderProfile, because a profile binds exactly one model (internal/config/types.go:40), so two profiles for the same model would each get a private bucket, which is not what anyone means by a per-model limit. Run ids throughmodelregistry.ResolveIDso aliases and deprecated names land in the same bucket.Enforcement: rpm only, in a wrapping
TurnSessionProvider, sliding window, in-process, pre-send. Token and cost ceilings ride on the existingOnUsageaccounting and gate the next request, documented as such rather than dressed up as real-time caps. ReuseCalculateCost; do not re-derive rates fromInputPerMillion.Pausing and prompting: fine in the TUI. In
zero execthere is nobody to ask, so it has to be a clean terminal error, with the message living ininternal/errhint/errhint.gonext to the existingRateLimitcase.The part that is a bad fit
A single budget shared across everything you run is not a small feature here, and I would keep it out of the first version.
zero execruns are separate OS processes, sub-agents are literally childzero execprocesses (internal/specialist/exec.go), and swarm members and cron runs each get their own independent retry schedule. An in-process limiter does not bound any of them together. Making it global means a durable counter plus a locking story across processes, which is a different and much larger change than what this issue asks for.Related, and worth knowing before someone designs a per-model cost budget that survives a restart: persisted usage events only carry a model id under
--allow-escalation(internal/cli/exec.go, in theOnUsagecallback), the TUI writes none, andusage.Reporthas no by-model dimension at all.BuildReportresolves a model purely to price a row and then discards it. So durable per-model attribution does not exist yet. Writing themodelkey unconditionally is a small prerequisite, and the reader already handles it.Next step
The decision blocking this is one question: is the first version per-process, or shared across processes? Everything else follows from that answer, and I would argue for per-process and documented as per-process.
If that is the answer, the first PR is small and self-contained: rpm only, as a
TurnSessionProviderwrapper installed whereOptimizedTurnSessionsalready installs one, plus the config key and its plumbing. That is reviewable on its own, and it is the piece that actually stops the runaway loop people are hitting. Token and cost ceilings can land after, on top of the accounting that already exists.Happy to review that if someone wants to pick it up.
Reacted by Chris ChenThanks for the unusually detailed map. I agree that the honest first slice is per-process RPM only: sliding-window, pre-send, at "TurnSessionProvider". TPM/cost limits and cross-process enforcement should remain explicit non-goals for that first step.
I maintain AEGIS Core, but its current implementation is Python and process-local, while Zero is Go. I do not want to imply that the Python SDK can simply be dropped into Zero or introduce an inappropriate runtime dependency.
Would you be open to an isolated fork experiment that implements the same narrow execution-boundary contract natively at "TurnSessionProvider": RPM only, per process, with a clean terminal error for "zero exec" and tests proving that requests above the configured window are blocked before provider I/O?
I would post the branch and measurements here before opening any PR, and I would not touch TPM, cost, or shared cross-process budgets. If that would be useful, could you confirm the preferred base branch and whether a draft PR would be welcome?
Yes, and that scope is exactly the one I would take: RPM only, per process, sliding window, pre-send, in a
TurnSessionProviderwrapper installed whereproviders.OptimizedTurnSessionsinstalls its own, withzero execfailing on a clean terminal error and the TUI free to do something friendlier later. TPM, cost and cross-process budgets stay out, and say so in the PR description so nobody reviews it against the larger ask.Base it on
main. A draft PR is welcome as soon as it compiles; you do not need to post measurements here first, the PR is the better place for them. What I will look for: the config key onFileConfigkeyed by resolved model id (run it throughmodelregistry.ResolveID), the limiter blocking insideStreambefore any provider I/O with a test proving request N+1 in the window never reaches the transport, the exec error text ininternal/errhintnext to the existing rate-limit case, and no new runtime dependency. Native Go only; no bridge to the Python implementation.One honest caveat: whether the feature ships is the owner's call (@kevincodex1), and this issue has had one contributor argue against the premise. A small, self-contained PR with the non-goals stated up front is the best way to make that decision easy. I will review it.
Thanks — I have now read
AGENTS.mdandCONTRIBUTING.md.I understand and will keep the scope exactly as described: RPM only, per process, sliding window, pre-send, implemented natively as a
TurnSessionProviderwrapper; no TPM, cost, cross-process budget, Python dependency, or broader refactor.Before starting implementation, I noticed that #390 currently has the
enhancementlabel but notissue-approved, which the contribution policy requires for community implementation and pull requests.If your invitation to open a draft PR is intended as formal approval for that narrow scope, could a maintainer please add the
issue-approvedlabel? I will wait for the label before changing code or opening the draft PR.- addedissue-approvedReviewed and approved by the core team; community PRs may implement this issue.Reviewed and approved by the core team; community PRs may implement this issue.
on Sep 24, 2026 @LortuArte sorry, this one slipped past me for over a week. Yes, it was meant as approval for exactly that scope, and I've added
issue-approvednow.Everything in my 12 September comment still stands: RPM only, per process, sliding window, pre-send in a
TurnSessionProviderwrapper, config onFileConfigkeyed by the resolved model id, and a cleanzero execerror. Base it onmainand open the draft whenever it compiles.
Idea: Per-model rate limiting and cost budgets
Problem
Zero runs with multiple model providers (OpenAI, Anthropic, Gemini, etc.) but has no built-in guardrails against runaway spending. A single misconfigured loop or prompt can generate thousands of API calls.
Proposed solution
Add configurable per-model rate limits and cost budgets in the project config:
When a limit is hit, zero pauses and prompts the user before continuing — giving visibility and control before a bill explodes.
This complements the existing network permission model by adding a cost layer.