Token Quotas and Rate Limiting
Quota Scopes
You can define quotas at four ownership levels. When more than one level applies to a request, the strictest (lowest) limit is the one enforced:
The tightest scope — a limit tied to a single API credential (LLM provider connection).
Shared across every credential that belongs to the same organization.
A credential's usage through a specific API proxy.
The broadest scope — all AI traffic under a project (the tenant boundary in multi-tenant setups).
Within any scope, you can also set model-specific limits — for example, a lower quota for a premium model and a higher one for a budget model. A model-specific limit overrides the scope's base limit only for requests that use that model; every other model keeps using the base limit.
Time Windows
Each scope can carry limits across four token windows, plus cost windows in USD:
A hard ceiling on tokens consumed in the current minute — useful for smoothing sudden traffic bursts.
The most commonly used window for everyday rate control, with its own optional USD budget.
A daily ceiling on token usage, with its own optional USD budget.
A monthly ceiling on token usage, with its own optional USD budget — the most common way to cap overall spend.
Token limits and USD budgets are independent — you can set only token limits, only a USD budget, or both together at any window.
Configuring a Quota
Go to AI Gateway → Token Quotas.
Choose Credential, Organization, Application, or Project.
Enter token limits for the minute, hour, day, and/or month windows. These apply to every model by default.
Select a model and enter a different set of limits for it. This overrides the base limit only for requests using that model.
Enter a USD amount for the hourly, daily, and/or monthly window.
Quotas take effect immediately for new requests.
Overflow Actions
When a request would exceed a quota, you choose what happens next:
The request is rejected outright.
The request is routed to the next provider or model in the failover chain instead of being rejected.
The request goes through as usual; only a warning is raised.
The request is automatically redirected to a lower-cost model instead of being blocked.
Per-Request Size Limit
Separate from the quota scopes above, you can cap the size of a single request — checked before the time-windowed quota, as a cheap pre-flight step. This lives in the same Token Rate Limit policy, under Oversized Guard:
| Field | Description | Default |
|---|---|---|
| Max Tokens Per Request | Blocks a single request whose estimated token count exceeds this value. Empty = no limit. | — |
| Max Prompt Characters | Blocks a single request whose prompt character count exceeds this value — only the prompt text itself is counted, not JSON envelope fields such as model or stream. Empty = no limit. | — |
| Limit Source | Where the effective token limit comes from: Explicit uses Max Tokens Per Request above; Model Catalog derives it from the selected model's context window instead (context window minus max output tokens minus a safety margin). If the model's context window is unknown, the limit is skipped for that request rather than blocking it. | Explicit |
| On Overflow | What happens when the token limit is exceeded: Block rejects the request; Truncate Oldest Messages drops the oldest conversation turns instead, until the request fits. Character-count overflow always blocks, regardless of this setting. | Block |
When a request is truncated, Apinizer always preserves every system message and the last user message — if those alone still exceed the limit, it falls back to Block rather than send a partial conversation. A successful truncation is tracked in three places: the X-Apinizer-AI-Truncated response header (number of dropped messages), the Guardrail Hits report (as truncated, alongside oversized), and — if AI Trace is active on the proxy — that request's trace detail.
This is a different concept from the quota-tier Overflow Actions above: quota overflow is about a scope's minute/hour/day/month budget running out over time; the per-request size limit is about a single request being too large for the model's context window, independent of any quota.
Threshold Alarms
You can raise alerts as usage approaches a quota, independent of the overflow action: at 50%, 80%, 90%, and 100% of the limit. These alerts appear alongside your other AI alerts so you can react before a quota actually blocks traffic.
How Quota Tracking Stays Accurate
Before a request is sent to the model, Apinizer estimates its token cost and reserves that amount against every applicable scope and window. Once the response finishes — including streamed responses — the reservation is adjusted to match the actual tokens consumed. This keeps concurrent requests from over-counting or under-counting against your quotas.
Monitoring Quota Usage
View real-time quota consumption in Reports and Analytics:
- Tokens Used and Tokens Remaining for the current window
- Cost to Date against any configured budget
- Alerts raised as usage approaches a limit
Token quotas can also be viewed and updated through the APIops REST API; see API Reference: AI Budgets.