Skip to content

Rate limits

View as Markdown

Limits keep the platform fast for everyone and protect you from runaway costs and abuse. When you reach one, you get a clear error with a stable code and, where waiting helps, a Retry-After header.

Who Limit When exceeded
Secret key 600 requests a minute per key 429 rate_limited
Publishable key (widget) 60 requests a minute per IP address, and 20 messages a minute per device 429 rate_limited
Dashboard login 10 attempts a minute per IP address, and per email 429 rate_limited
Sign-up 5 a day per IP address 429 rate_limited
What Limit When exceeded
Messages waiting in one session (queue) 10 409 session_queue_full
A message while the session is answering (reject) — 409 session_busy, Retry-After: 2
Anonymous widget sessions 20 an hour per IP address 429 anonymous_limit_reached and a polite notice
Anonymous widget messages 60 an hour per IP address; 40 per session 429 anonymous_limit_reached and a polite notice
Anonymous conversations per agent widget.anonymous_daily_conversations a day (default 200) A polite “try again later” notice
Test-panel runs on platform models 200 a day per project (unlimited with your own provider key) 429 playground_limit_reached
AI conversations on the Free plan 100 a month Sessions hand off with a notice; ask gets 429 quota_exceeded
Daily model cost per organization Free $1, Starter $10, Business $30, Scale $100 Sessions hand off with a notice; ask gets 429 cost_cap_exceeded
What Limit
Tool calls executed per round 5 (extra calls get too_many_calls)
Model-and-tool rounds per turn tools.max_tool_rounds, 1–5 (default 3)
HTTP tool calls per end user 20 an hour across all HTTP tools, counting every attempt; per tool limits.per_end_user_per_hour (1–100)
create_ticket per end user 5 a day
HTTP tool time up to 15 seconds (connect 3 seconds)
HTTP tool response up to 1 MiB; at most 20 rows and 12 fields reach the model, within 4 KiB
Client tool outputs due within 10 minutes

When a tool limit is reached, the model is told in neutral words and can answer from what it has or hand the conversation to your team; your API never sees the request.

What Limit When exceeded
input with a secret key 32,000 characters 413 input_too_large
input with a publishable key or client token 4,000 characters 413 input_too_large
Request body 1 MiB (5 MiB for knowledge sources) 413 request_too_large
messages on the OpenAI-compatible endpoint 100 Rejected with an OpenAI-style error
Page size of lists 100 (default 20) Capped
metadata 16 keys; keys up to 64 characters, values up to 512 422 validation_failed
Idempotency-Key 255 characters 400 invalid_idempotency_key
Webhook endpoints per project 10 422 webhook_endpoint_limit_reached
Header When Meaning
Retry-After 429 and some 409 responses Seconds to wait before retrying.
RateLimit-Policy Rate-limited responses The limit that applies: its quota and window.
RateLimit Rate-limited responses What remains in the current window and when it resets.
x-should-retry Errors where retrying cannot help false — for example quota_exceeded.

RateLimit-Policy and RateLimit follow the IETF HTTP RateLimit header fields draft.

  • Honor Retry-After. Wait at least that long; add a little random jitter so many clients don’t retry at the same instant.
  • Back off exponentially for repeated 429 and 5xx responses: for example 0.5 s, 1 s, 2 s, 4 s, then give up and report.
  • Don’t retry when x-should-retry is false.
  • Make retries safe with an Idempotency-Key, so a retry never creates a duplicate.
  • Queue, don’t burst. If you import or migrate in bulk, run a fixed number of workers instead of firing every request at once.
  • Identify signed-in users in the widget with client tokens: verified users aren’t subject to the anonymous limits.

Need higher limits? Contact us — limits are set per plan and can be raised for Enterprise agreements.