Skip to content

Per-model rate limiting and cost budgets #390

Description

@fuleinist

Idea: Per-model rate limiting and cost budgets

Problem

Zero runs with multiple model providers (OpenAI, Anthropic, Gemini, etc.) but has no built-in guardrails against runaway spending. A single misconfigured loop or prompt can generate thousands of API calls.

Proposed solution

Add configurable per-model rate limits and cost budgets in the project config:

models:
  - name: gpt-4o
    rpm: 60        # requests per minute
    tpm: 100000    # tokens per minute
    max_cost: 10   # USD per session

When a limit is hit, zero pauses and prompts the user before continuing — giving visibility and control before a bill explodes.

This complements the existing network permission model by adding a cost layer.

Activity

  1. KunjShah95 commented on Jul 3, 2026

    @KunjShah95
    Contributor

    Shall I work on this features

  2. lfaoro commented on Jul 15, 2026

    @lfaoro
    Contributor

    you should control your token spend limits directly via the provider
    all the providers I used have a limit setting

    the token spend calculations a harness does might not reflect exactly the provider costs as prices change, extra costs could be added, cache misses, tool calls, etc..

  3. LortuArte commented on Aug 6, 2026

    @LortuArte

    Correction — 24 September 2026. I am the author/maintainer of AEGIS.

    I withdraw the claims about physically killing threads, <0.005 ms full processing, and replacing Zero's shared or distributed rate/budget controls. AEGIS is Python; the later rpm-only Go experiment discussed in this issue is separate and is not evidence of AEGIS adoption.

    AEGIS Core 3.4.0 provides a Python process-local authorization gate. The caller must supply a known amount, reuse the logical call ID and execute the tool only when execution_permitted is true. A cached authorization receipt is not the tool's result. This does not establish multiprocess coordination, crash recovery or exactly-once external effects. Current scope: https://github.com/LortuArte/aegis-sdk.

  4. Vasanthdev2004 commented on Sep 4, 2026

    @Vasanthdev2004
    Collaborator

    Short answer: no, Zero cannot do any of this today. There is no rpm limit, no tpm limit, and no cost cap anywhere in the tree. I grepped for the obvious spellings (rate.Limiter, x/time/rate, RequestsPerMinute, TokensPerMinute, CostLimit, MaxCost, SpendLimit, budgetUSD) and got zero hits. Not a config key, not a flag, not an env var.

    Sorry this sat so long. Here is what is actually there, because a good chunk of the work is already done and it is not the part people expect.

    The accounting layer is real and finished

    This is the good news. zeroruntime.Usage (internal/zeroruntime/types.go:158) is a normalized per-request record that all three first-party providers emit, with a documented invariant worth reading before you touch it: cache-read and cache-write are subsets of InputTokens, not additive, and ReasoningTokens is a subset of OutputTokens.

    modelregistry.CalculateCost (internal/modelregistry/cost.go:38) already turns one of those into a USD breakdown including tiered pricing and the cache splits. usage.Tracker (internal/usage/tracker.go) already aggregates it per model, and Summary().ByModel already answers "N tokens and $X on model M".

    So nobody needs to build counting or pricing. One catch: the tracker is constructed only at internal/tui/model.go:912 and internal/tui/btw.go:112. It is TUI-only and in-memory. Headless zero exec never builds one at all.

    Exactly one thing enforces a budget today

    Goal.TokenBudget, via Store.AddGoalUsage (internal/sessions/goal.go:280). It accumulates tokens, compares against the budget, flips the goal to GoalStatusBudgetLimited with reason "token budget reached", and emits an event. That is the shipped shape of "a budget that stops work" in this codebase, and a new budget should mirror it rather than inventing a second state machine.

    But look at where it is called from. One non-test caller: internal/tui/goal.go:325, inside reconcileGoalAfterRun, which runs after a whole run has finished and sums the run's usage events. So it cannot interrupt anything. A single run overshoots the budget by however much it wants, and the budget only gates whether the next autonomous continuation starts. It is also TUI-only, so exec, ACP and cron never account it at all.

    Which is a concrete instance of exactly the failure the thread already flagged.

    The enforcement point

    The warning upthread is correct, and I would put it more strongly: there is nothing to make a network call to. Any limit Zero enforces is a local counter, so the question is not "which service do we ask" but "which local seam sees every request before it goes out".

    That seam exists and is purpose-built. zeroruntime.TurnSession (internal/zeroruntime/session.go:47) is opened once per run, and the loop reassigns its provider local so everything flows through it. From internal/agent/loop.go at the assignment:

    Route ALL provider I/O, per-turn streams, mid-stream retries, the compaction summarizer, and the final-answer request, through the session by reassigning the loop's provider local. Every existing call site reads this local, so no signatures change anywhere downstream.

    That is the place. A limiting TurnSessionProvider wrapping zeroruntime.NewProviderTurnSessionProvider needs no signature change anywhere, and providers.OptimizedTurnSessions (internal/providers/turn_session.go:50) is the existing example of installing one behind a gate. Blocking inside Stream is in-process and pre-send, so a runaway loop stops on request N, not request N plus whatever the round trip let through.

    The thing not to use is agent.Options.OnUsage (internal/agent/types.go:417). It fires for every request including the compaction summarizer, which makes it the right place to account, but it is func(Usage) with no return and no error. It cannot refuse anything. Wire accounting there, enforcement in the session.

    There is a lower option, providerio.SendWithAuthRetry (internal/providers/providerio/auth.go:31), where the SpanProviderQueue block already carries this comment: "No send-side semaphore exists today; a future request queue would also accumulate here." That is provider-agnostic and gets tracing for free, but it does not know token counts, so it only really serves rpm.

    What Zero can and cannot honestly know

    This is where I would push back on parts of the ask, because the three limits are not equally knowable.

    rpm is exact. A request count is known locally and precisely before the send. This is the one that works properly, and it is also the one that actually stops a runaway loop.

    tpm is approximate going in and unknown coming out. The only token estimator in the repo is estimateTokens (internal/agent/compaction.go:124), and its own doc says it "deliberately uses no real tokenizer; it only needs to be monotonic and roughly proportional". It is characters over four. Output tokens are not known until the response arrives. So a tpm limit can only honestly be "refuse the next request because the trailing window is already over", never "hold the window under N".

    cost is a reconstructed estimate, and sometimes it is fiction. Rates are hardcoded in internal/modelregistry/catalog.go with sourceLastVerified = "2026-06-04", optionally overlaid from models.dev. Unpriced or unknown models contribute tokens and no cost, by design. And for anyone using a ChatGPT or Claude subscription through the local-proxy recipe in docs/oauth-subscriptions.md, a per-token dollar figure does not correspond to anything they are billed. A cost cap has to say out loud that it is an estimate, the way zero usage report already does.

    Smallest design that is honest

    Config: a new top-level limits key on FileConfig (internal/config/types.go:352), keyed by resolved model id. Not on ProviderProfile, because a profile binds exactly one model (internal/config/types.go:40), so two profiles for the same model would each get a private bucket, which is not what anyone means by a per-model limit. Run ids through modelregistry.ResolveID so aliases and deprecated names land in the same bucket.

    Enforcement: rpm only, in a wrapping TurnSessionProvider, sliding window, in-process, pre-send. Token and cost ceilings ride on the existing OnUsage accounting and gate the next request, documented as such rather than dressed up as real-time caps. Reuse CalculateCost; do not re-derive rates from InputPerMillion.

    Pausing and prompting: fine in the TUI. In zero exec there is nobody to ask, so it has to be a clean terminal error, with the message living in internal/errhint/errhint.go next to the existing RateLimit case.

    The part that is a bad fit

    A single budget shared across everything you run is not a small feature here, and I would keep it out of the first version. zero exec runs are separate OS processes, sub-agents are literally child zero exec processes (internal/specialist/exec.go), and swarm members and cron runs each get their own independent retry schedule. An in-process limiter does not bound any of them together. Making it global means a durable counter plus a locking story across processes, which is a different and much larger change than what this issue asks for.

    Related, and worth knowing before someone designs a per-model cost budget that survives a restart: persisted usage events only carry a model id under --allow-escalation (internal/cli/exec.go, in the OnUsage callback), the TUI writes none, and usage.Report has no by-model dimension at all. BuildReport resolves a model purely to price a row and then discards it. So durable per-model attribution does not exist yet. Writing the model key unconditionally is a small prerequisite, and the reader already handles it.

    Next step

    The decision blocking this is one question: is the first version per-process, or shared across processes? Everything else follows from that answer, and I would argue for per-process and documented as per-process.

    If that is the answer, the first PR is small and self-contained: rpm only, as a TurnSessionProvider wrapper installed where OptimizedTurnSessions already installs one, plus the config key and its plumbing. That is reviewable on its own, and it is the piece that actually stops the runaway loop people are hitting. Token and cost ceilings can land after, on top of the accounting that already exists.

    Happy to review that if someone wants to pick it up.

  5. LortuArte commented on Sep 9, 2026

    @LortuArte

    Thanks for the unusually detailed map. I agree that the honest first slice is per-process RPM only: sliding-window, pre-send, at "TurnSessionProvider". TPM/cost limits and cross-process enforcement should remain explicit non-goals for that first step.

    I maintain AEGIS Core, but its current implementation is Python and process-local, while Zero is Go. I do not want to imply that the Python SDK can simply be dropped into Zero or introduce an inappropriate runtime dependency.

    Would you be open to an isolated fork experiment that implements the same narrow execution-boundary contract natively at "TurnSessionProvider": RPM only, per process, with a clean terminal error for "zero exec" and tests proving that requests above the configured window are blocked before provider I/O?

    I would post the branch and measurements here before opening any PR, and I would not touch TPM, cost, or shared cross-process budgets. If that would be useful, could you confirm the preferred base branch and whether a draft PR would be welcome?

  6. Vasanthdev2004 commented on Sep 12, 2026

    @Vasanthdev2004
    Collaborator

    Yes, and that scope is exactly the one I would take: RPM only, per process, sliding window, pre-send, in a TurnSessionProvider wrapper installed where providers.OptimizedTurnSessions installs its own, with zero exec failing on a clean terminal error and the TUI free to do something friendlier later. TPM, cost and cross-process budgets stay out, and say so in the PR description so nobody reviews it against the larger ask.

    Base it on main. A draft PR is welcome as soon as it compiles; you do not need to post measurements here first, the PR is the better place for them. What I will look for: the config key on FileConfig keyed by resolved model id (run it through modelregistry.ResolveID), the limiter blocking inside Stream before any provider I/O with a test proving request N+1 in the window never reaches the transport, the exec error text in internal/errhint next to the existing rate-limit case, and no new runtime dependency. Native Go only; no bridge to the Python implementation.

    One honest caveat: whether the feature ships is the owner's call (@kevincodex1), and this issue has had one contributor argue against the premise. A small, self-contained PR with the non-goals stated up front is the best way to make that decision easy. I will review it.

  7. LortuArte commented on Sep 13, 2026

    @LortuArte

    Thanks — I have now read AGENTS.md and CONTRIBUTING.md.

    I understand and will keep the scope exactly as described: RPM only, per process, sliding window, pre-send, implemented natively as a TurnSessionProvider wrapper; no TPM, cost, cross-process budget, Python dependency, or broader refactor.

    Before starting implementation, I noticed that #390 currently has the enhancement label but not issue-approved, which the contribution policy requires for community implementation and pull requests.

    If your invitation to open a draft PR is intended as formal approval for that narrow scope, could a maintainer please add the issue-approved label? I will wait for the label before changing code or opening the draft PR.

  8. added
    issue-approvedReviewed and approved by the core team; community PRs may implement this issue.
    on Sep 24, 2026
  9. Vasanthdev2004 commented on Sep 24, 2026

    @Vasanthdev2004
    Collaborator

    @LortuArte sorry, this one slipped past me for over a week. Yes, it was meant as approval for exactly that scope, and I've added issue-approved now.

    Everything in my 12 September comment still stands: RPM only, per process, sliding window, pre-send in a TurnSessionProvider wrapper, config on FileConfig keyed by the resolved model id, and a clean zero exec error. Base it on main and open the draft whenever it compiles.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

enhancementNew feature or requestissue-approvedReviewed and approved by the core team; community PRs may implement this issue.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions