Rate limits
View as Markdown# Rate limits
> Request rates, conversation limits, size limits and the headers K-Agent sends when you reach one — and how to handle them.
Limits keep the platform fast for everyone and protect you from runaway costs and abuse. When you reach one, you get a clear error with a stable code and, where waiting helps, a `Retry-After` header.
## Request rates
| Who | Limit | When exceeded |
|---|---|---|
| Secret key | 600 requests a minute per key | `429 rate_limited` |
| Publishable key (widget) | 60 requests a minute per IP address, and 20 messages a minute per device | `429 rate_limited` |
| Dashboard login | 10 attempts a minute per IP address, and per email | `429 rate_limited` |
| Sign-up | 5 a day per IP address | `429 rate_limited` |
## Conversation limits
| What | Limit | When exceeded |
|---|---|---|
| Messages waiting in one session (`queue`) | 10 | `409 session_queue_full` |
| A message while the session is answering (`reject`) | — | `409 session_busy`, `Retry-After: 2` |
| Anonymous widget sessions | 20 an hour per IP address | `429 anonymous_limit_reached` and a polite notice |
| Anonymous widget messages | 60 an hour per IP address; 40 per session | `429 anonymous_limit_reached` and a polite notice |
| Anonymous conversations per agent | `widget.anonymous_daily_conversations` a day (default 200) | A polite "try again later" notice |
| Test-panel runs on platform models | 200 a day per project (unlimited with your own provider key) | `429 playground_limit_reached` |
| AI conversations on the Free plan | 100 a month | Sessions hand off with a notice; `ask` gets `429 quota_exceeded` |
| Daily model cost per organization | Free $1, Starter $10, Business $30, Scale $100 | Sessions hand off with a notice; `ask` gets `429 cost_cap_exceeded` |
## Tool limits
| What | Limit |
|---|---|
| Tool calls executed per round | 5 (extra calls get `too_many_calls`) |
| Model-and-tool rounds per turn | `tools.max_tool_rounds`, 1–5 (default 3) |
| HTTP tool calls per end user | 20 an hour across all HTTP tools, counting every attempt; per tool `limits.per_end_user_per_hour` (1–100) |
| `create_ticket` per end user | 5 a day |
| HTTP tool time | up to 15 seconds (connect 3 seconds) |
| HTTP tool response | up to 1 MiB; at most 20 rows and 12 fields reach the model, within 4 KiB |
| Client tool outputs | due within 10 minutes |
When a tool limit is reached, the model is told in neutral words and can answer from what it has or hand the conversation to your team; your API never sees the request.
## Size limits
| What | Limit | When exceeded |
|---|---|---|
| `input` with a secret key | 32,000 characters | `413 input_too_large` |
| `input` with a publishable key or client token | 4,000 characters | `413 input_too_large` |
| Request body | 1 MiB (5 MiB for knowledge sources) | `413 request_too_large` |
| `messages` on the OpenAI-compatible endpoint | 100 | Rejected with an OpenAI-style error |
| Page size of lists | 100 (default 20) | Capped |
| `metadata` | 16 keys; keys up to 64 characters, values up to 512 | `422 validation_failed` |
| `Idempotency-Key` | 255 characters | `400 invalid_idempotency_key` |
| Webhook endpoints per project | 10 | `422 webhook_endpoint_limit_reached` |
## Headers
| Header | When | Meaning |
|---|---|---|
| `Retry-After` | `429` and some `409` responses | Seconds to wait before retrying. |
| `RateLimit-Policy` | Rate-limited responses | The limit that applies: its quota and window. |
| `RateLimit` | Rate-limited responses | What remains in the current window and when it resets. |
| `x-should-retry` | Errors where retrying cannot help | `false` — for example `quota_exceeded`. |
`RateLimit-Policy` and `RateLimit` follow the IETF HTTP RateLimit header fields draft.
## Handling limits well
- **Honor `Retry-After`.** Wait at least that long; add a little random jitter so many clients don't retry at the same instant.
- **Back off exponentially** for repeated `429` and `5xx` responses: for example 0.5 s, 1 s, 2 s, 4 s, then give up and report.
- **Don't retry** when `x-should-retry` is `false`.
- **Make retries safe** with an `Idempotency-Key`, so a retry never creates a duplicate.
- **Queue, don't burst.** If you import or migrate in bulk, run a fixed number of workers instead of firing every request at once.
- **Identify signed-in users** in the widget with client tokens: verified users aren't subject to the anonymous limits.
Need higher limits? Contact us — limits are set per plan and can be raised for Enterprise agreements.
Limits keep the platform fast for everyone and protect you from runaway costs and abuse. When you reach one, you get a clear error with a stable code and, where waiting helps, a Retry-After header.
Request rates
Section titled “Request rates”| Who | Limit | When exceeded |
|---|---|---|
| Secret key | 600 requests a minute per key | 429 rate_limited |
| Publishable key (widget) | 60 requests a minute per IP address, and 20 messages a minute per device | 429 rate_limited |
| Dashboard login | 10 attempts a minute per IP address, and per email | 429 rate_limited |
| Sign-up | 5 a day per IP address | 429 rate_limited |
Conversation limits
Section titled “Conversation limits”| What | Limit | When exceeded |
|---|---|---|
Messages waiting in one session (queue) |
10 | 409 session_queue_full |
A message while the session is answering (reject) |
— | 409 session_busy, Retry-After: 2 |
| Anonymous widget sessions | 20 an hour per IP address | 429 anonymous_limit_reached and a polite notice |
| Anonymous widget messages | 60 an hour per IP address; 40 per session | 429 anonymous_limit_reached and a polite notice |
| Anonymous conversations per agent | widget.anonymous_daily_conversations a day (default 200) |
A polite “try again later” notice |
| Test-panel runs on platform models | 200 a day per project (unlimited with your own provider key) | 429 playground_limit_reached |
| AI conversations on the Free plan | 100 a month | Sessions hand off with a notice; ask gets 429 quota_exceeded |
| Daily model cost per organization | Free $1, Starter $10, Business $30, Scale $100 | Sessions hand off with a notice; ask gets 429 cost_cap_exceeded |
Tool limits
Section titled “Tool limits”| What | Limit |
|---|---|
| Tool calls executed per round | 5 (extra calls get too_many_calls) |
| Model-and-tool rounds per turn | tools.max_tool_rounds, 1–5 (default 3) |
| HTTP tool calls per end user | 20 an hour across all HTTP tools, counting every attempt; per tool limits.per_end_user_per_hour (1–100) |
create_ticket per end user |
5 a day |
| HTTP tool time | up to 15 seconds (connect 3 seconds) |
| HTTP tool response | up to 1 MiB; at most 20 rows and 12 fields reach the model, within 4 KiB |
| Client tool outputs | due within 10 minutes |
When a tool limit is reached, the model is told in neutral words and can answer from what it has or hand the conversation to your team; your API never sees the request.
Size limits
Section titled “Size limits”| What | Limit | When exceeded |
|---|---|---|
input with a secret key |
32,000 characters | 413 input_too_large |
input with a publishable key or client token |
4,000 characters | 413 input_too_large |
| Request body | 1 MiB (5 MiB for knowledge sources) | 413 request_too_large |
messages on the OpenAI-compatible endpoint |
100 | Rejected with an OpenAI-style error |
| Page size of lists | 100 (default 20) | Capped |
metadata |
16 keys; keys up to 64 characters, values up to 512 | 422 validation_failed |
Idempotency-Key |
255 characters | 400 invalid_idempotency_key |
| Webhook endpoints per project | 10 | 422 webhook_endpoint_limit_reached |
Headers
Section titled “Headers”| Header | When | Meaning |
|---|---|---|
Retry-After |
429 and some 409 responses |
Seconds to wait before retrying. |
RateLimit-Policy |
Rate-limited responses | The limit that applies: its quota and window. |
RateLimit |
Rate-limited responses | What remains in the current window and when it resets. |
x-should-retry |
Errors where retrying cannot help | false — for example quota_exceeded. |
RateLimit-Policy and RateLimit follow the IETF HTTP RateLimit header fields draft.
Handling limits well
Section titled “Handling limits well”- Honor
Retry-After. Wait at least that long; add a little random jitter so many clients don’t retry at the same instant. - Back off exponentially for repeated
429and5xxresponses: for example 0.5 s, 1 s, 2 s, 4 s, then give up and report. - Don’t retry when
x-should-retryisfalse. - Make retries safe with an
Idempotency-Key, so a retry never creates a duplicate. - Queue, don’t burst. If you import or migrate in bulk, run a fixed number of workers instead of firing every request at once.
- Identify signed-in users in the widget with client tokens: verified users aren’t subject to the anonymous limits.
Need higher limits? Contact us — limits are set per plan and can be raised for Enterprise agreements.