# Rate limits and quotas

Every limit, what it is keyed on, the headers that report it, and how to back off.

Limits keep one runaway loop from hurting everyone else. They are never a CAPTCHA: every limited response says how long to wait, in headers an agent can read.

## Request rate limits

Fixed windows

| What | Limit | Counted per |
|---|---|---|
| Agent endpoints (`/v1/execute`, `/v1/recommend`, receipts, jobs, disputes, approvals, `/mcp`) | 600 requests a minute | agent key |
| Org endpoints (`/v1/orgs/{orgId}/…`) | 300 requests a minute | org key or signed-in session (IP when neither) |
| Public data (`/v1/tools`, `/v1/scores`, `/v1/pricing`, …) | 120 requests a minute | IP address |
| Sign-in links (`POST /auth/email/start`) | 20 an hour | IP address (and at most 5 emails an hour to one address) |
| Waitlist (`POST /v1/waitlist`) | 10 an hour | IP address |

## Headers

Every limited endpoint answers with these headers, on success too, so you can pace yourself before you hit the wall:

Rate-limit headers

| Header | Value |
|---|---|
| `RateLimit-Limit` | Requests allowed in the window. |
| `RateLimit-Remaining` | Requests left in this window. |
| `RateLimit-Reset` | Seconds until the window resets. |
| `RateLimit-Policy` | The limit and window, e.g. `600;w=60`. |
| `Retry-After` | Only on a 429: seconds to wait before trying again. |

**429 response**

```http
HTTP/1.1 429 Too Many Requests
RateLimit-Limit: 600
RateLimit-Remaining: 0
RateLimit-Reset: 17
RateLimit-Policy: 600;w=60
Retry-After: 17

{ "error": { "code": "rate_limited", "message": "Too many requests from this agent. Try again in 17s.", "doc_url": "…/docs/errors#rate_limited", "retryable": true } }
```

## Daily quotas

`POST /v1/recommend` counts against a daily quota per org, by plan: every call is a score lookup, and a call with free text (`task`) instead of a `task_type` also counts one free-text lookup. Over the quota: `429 plan_limit`. Quotas reset at 00:00 UTC.

Recommend quotas per day

| Plan | Score lookups | Free-text lookups |
|---|---|---|
| Pay as you go | 1,000 | 100 |
| Pro | 10,000 | 500 |
| Max | no fixed limit | no fixed limit |

- **Provider quotas.** Some providers cap how many calls one customer may make in a day. Past it, requests to that provider's tools get `429 provider_quota` (reset at 00:00 UTC); other tools for the same task keep working.
- **People lookups.** At most 500 people lookups per company domain per day: `429 aup_limit`. See [data handling](/docs/data-handling#acceptable-use).
- **Trial spend.** Until an org's first top-up, at most 50 trial credits an hour: `429 trial_limit`.
- **Open disputes.** At most 100 open at once per org.

## Size limits

Size limits

| What | Limit | Over it |
|---|---|---|
| Request body | 1 MB | `413 payload_too_large` |
| Inputs in one `execute` | 1,000 | `400 invalid_input` |
| Items run straight away | 25; more runs as a [job](/docs/jobs) | — |
| Job run time | 2 hours; items not yet run are released as `expired` | — |
| Webhook endpoints per org | 10 | `402 plan_limit` |

## Backing off

- On a 429, wait `Retry-After` seconds, then retry. Only codes with `retryable: true` are worth retrying ([error codes](/docs/errors)).
- Without `Retry-After` (a 5xx or a network error), back off exponentially with jitter.
- Retrying `execute` is safe only with the same `idempotency_key`: the retry returns the first answer and is never charged twice.
- The SDKs do all of this for you: up to 3 retries by default, waiting out `Retry-After` up to 60 seconds, otherwise exponential backoff from 0.5 s up to 8 s with jitter, and the same idempotency key on every `execute` retry. See [SDKs](/docs/sdks#retries).

Updated 2026-09-30. This page as HTML: https://arettic.com/docs/rate-limits · Markdown: https://arettic.com/docs/rate-limits.md · JSON: https://arettic.com/docs/rate-limits.json
