# Pass rules

Six task types, one published check each: what passes, what fails, and what is refunded.

Every result bought through Arettic is checked before you pay for it. The check is a pass rule: one per task type, published here, versioned, and run on every live and test result the same way. A pass is charged. A fail is refunded in full and its data is withheld. Only `web_search` can be partial, and a partial is charged pro rata. This page has the rules, the free input check that runs before any provider is called, what each rule does to a result field by field, the reason codes, how versions change, and how to read the rules from the API.

> Arettic is pre-launch. Keys go to design partners first; everyone else joins the [waitlist](/waitlist). The `@arettic/sdk` and `@arettic/mcp` packages (npm) and the `arettic` package (PyPI) are published at launch. A test key (`sk_test_…`) runs these same rules on mock results and charges nothing; see [Test mode](/docs/test-mode).

## The rules

Each rule is named `task_type@version`, for example `find_email@v1`. That name is the `pass_rule` on every tool in [recommend](/docs/recommend) and `GET /v1/tools`, and its version is the `check_version` on every execute answer and every receipt. This table is generated from the same published constants the API serves at `GET /v1/task-types` and `GET /v1/formula`, so it cannot drift from what the API does.

The published pass rules

| task_type | Used for | Version | Passes if | Partial or failed |
|---|---|---|---|---|
| find_email | Outbound prospecting, recruiting outreach, partner sourcing | v1 | An email is returned and its verification status is valid (not catch-all or unknown) | Catch-all counts as a fail and is refunded |
| verify_email | Cleaning a list before a campaign, sign-up and CRM hygiene | v1 | The verifier returns a definitive status (valid or invalid) | Unknown or timeout counts as a fail |
| enrich_company | Account research, lead scoring and routing, CRM enrichment | v1 | The record matches the requested domain (or name + country) and has name, domain, employee range and industry | Missing required fields counts as a fail |
| enrich_person | Lead qualification, contact research, candidate and investor research | v1 | Name (fuzzy ≥ 0.9) and company domain (exact) match the input; a title and one contact field are present | A match without a contact field counts as a fail |
| web_search | Market and competitor research, news monitoring, building target lists | v1 | At least N results (5 by default) with valid, non-duplicate URLs | Fewer than N is charged pro rata |
| extract_url | Reading pricing pages, docs, job posts and filings into an agent | v1 | HTTP 200 with at least 200 characters of main content, not a block or captcha page | A blocked page counts as a fail |

The rule is a pure function of your input and the provider's result. The same result always gets the same verdict, whichever provider produced it and whoever bought it. Nothing about a tool's score, price or provider changes the verdict.

## Where the check sits in a purchase

1. **Input check, free.** Every input is checked for syntax, then domains are looked up in DNS and URLs are checked for a public host. A problem answers [invalid_input](/docs/errors#invalid_input) (HTTP 400) and no provider is called. See [Before the call](#before-the-call).
2. **Hold.** The tool's price for your plan is held per item.
3. **Call.** Arettic calls the provider with its own credentials. Each attempt is cut off after 30 seconds, with one retry after an error or a timeout.
4. **Check.** The rule for the tool's task type runs on the answer and gives `pass`, `partial` or `fail`, with a reason for anything but a pass. For `find_email`, Arettic first re-verifies the address with its own verifier and the rule judges that verdict. See [Our own verifier](#our-own-verifier).
5. **Settle.** Straight after the check. A pass captures the price. A partial captures `price × fraction`, rounded up to a whole credit, and releases the rest. A fail releases everything. A provider error, a timeout or an error in the checker itself also releases everything. See [Never charged](#never-charged).
6. **Deliver.** A passed or partial item comes back with its `result`. A failed item comes back with its `reason` only: the paid data is withheld.

The full request and response, every status and every reason code are on [Execute](/docs/execute). A charge you believe was wrong can be disputed within the window on [Disputes](/docs/disputes); the reviewer sees the check, its version and the reason it recorded.

## Before the call: the free input check

A request that fails the input check never reaches a provider and costs nothing. Two checks run in order: the syntax check on every item, and only if every item passes it, the network check. Every problem is reported at once, the first ten in the message, in the form `input.field: problem` for one input or `inputs[3].field: problem` for a batch, followed by `; and N more` when there are more than ten, and ending with `Nothing was charged.` An input that is not a JSON object is reported as `input: must be an object`.

### The syntax check

This is what each task type needs to get past the syntax check, and the exact problem text you get when it does not. A domain is normalised before it is judged: trimmed, lower-cased, and stripped of a leading `http://` or `https://`, a leading `www.` and anything from the first `/`, `?` or `#`. What is left must be 3 to 253 characters of labels made of letters, digits and hyphens (no hyphen at a label's start or end), ending in a top-level label of 2 to 63 letters. So `https://www.acme.com/about` is accepted as `acme.com`.

The syntax check per task type

| task_type | What is accepted | Problem text when it is not |
|---|---|---|
| find_email | `first_name` and `last_name`: non-empty strings of up to 100 characters. `domain`: a domain like `acme.com`. | `first_name: required`, `last_name: required`, `domain: must be a domain like acme.com` |
| verify_email | `email`: an address of up to 254 characters with a local part, an `@` and a domain whose last label has at least 2 characters. | `email: must be an email address` |
| enrich_company | `domain`: a domain, or instead `name`: a non-empty string of up to 200 characters. If neither is usable the problem is reported on `domain`. | `domain: send a domain (or a company name, optionally with country)` |
| enrich_person | `first_name` and `last_name`: non-empty strings of up to 100 characters. `company_domain`: a domain. | `first_name: required`, `last_name: required`, `company_domain: must be a domain like acme.com` |
| web_search | `query`: a non-empty string of up to 500 characters. `n`: absent, or a whole number from 1 to 25. | `query: required, up to 500 characters`, `n: must be a whole number from 1 to 25` |
| extract_url | `url`: a string that parses as a full URL, with the scheme `http` or `https`. | `url: required`, `url: must be a full URL`, `url: must be http or https` |

### The network check

With a live key, once every item passes the syntax check, the domains in the request are looked up in DNS and the URLs are checked for a public host. A test key skips this step, because mock tools use made-up domains.

- **Domains must exist.** For `find_email` and `enrich_company` the `domain`, for `enrich_person` the `company_domain`, and for `verify_email` the part of `email` after the `@` are each looked up for MX, A and AAAA records. Only a definite answer that the domain does not exist rejects the item, with the problem `the domain acme.example doesn't exist`. If Arettic's own DNS lookup fails or takes more than 2 seconds, the item goes through: an outage on Arettic's side never blocks your request. Answers are cached for an hour when the domain exists and for 10 minutes when it does not.
- **URLs must be public.** For `extract_url`, the host of `url` may not be `localhost`, end in `.localhost`, `.local` or `.internal`, or be a loopback, private, link-local or unspecified IP address. The problem is `url: must be a public web address`.

**Response (example): a batch with two bad inputs, HTTP 400**

```json
{
  "error": {
    "code": "invalid_input",
    "message": "inputs[1].domain: must be a domain like acme.com; inputs[2].first_name: required. Nothing was charged."
  }
}
```

## Each rule, field by field

For every task type: the input fields the tool takes, the output fields a passed `result` has, what the rule does to the result in order, and the reason it records when it stops. The first reason that applies wins. `no_result` always comes first: it means the provider had no record or answered with nothing usable. The input and output tables are generated from the published constants.

### find_email

**Passes if:** An email is returned and its verification status is valid (not catch-all or unknown).

find_email input

| Input field | Meaning |
|---|---|
| first_name | string |
| last_name | string |
| domain | company domain, e.g. acme.com |

find_email output

| Output field | Meaning |
|---|---|
| email | string |
| verification_status | valid | invalid | catch_all | unknown |
| confidence | 0–1, optional |

1. No result object, or no non-empty `email` in it: `no_result`.
2. `verification_status` is `catch_all`: `catch_all`. The domain accepts any address, so the one found cannot be confirmed. This is a fail and is refunded.
3. `verification_status` is anything but `valid`: `status:<value>`, for example `status:unknown` or `status:invalid`. A missing status counts as `status:unknown`.
4. Otherwise: pass.

With a live key, `verification_status` is Arettic's own verifier's verdict, not the finder's. The finder's own claim is replaced before the rule runs. See [Our own verifier](#our-own-verifier). Provider status names are normalised to the four published values, so `accept_all` from a provider is `catch_all` here. `confidence` is the finder's own score, from 0 to 1, when it gives one; the rule does not use it.

### verify_email

**Passes if:** The verifier returns a definitive status (valid or invalid).

verify_email input

| Input field | Meaning |
|---|---|
| email | string |

verify_email output

| Output field | Meaning |
|---|---|
| email | string |
| status | valid | invalid | catch_all | unknown |
| sub_status | optional |

1. No result object: `no_result`.
2. `status` is `valid` or `invalid`: pass. Both are definitive answers, and an invalid address is a true, useful result.
3. Any other `status`: `status:<value>`, for example `status:unknown` or `status:catch_all`. A missing status counts as `status:unknown`.

### enrich_company

**Passes if:** The record matches the requested domain (or name + country) and has name, domain, employee range and industry.

enrich_company input

| Input field | Meaning |
|---|---|
| domain | company domain |
| name | company name (if no domain) |
| country | optional, with name |

enrich_company output

| Output field | Meaning |
|---|---|
| name | string |
| domain | string |
| employee_range | e.g. 51-200 |
| industry | string |
| country | optional |

1. No result object, or an empty one: `no_result`.
2. You sent a `domain` and the record's `domain` is not the same after normalisation (lower-case, no scheme, no `www.`, no path): `domain_mismatch`.
3. You sent no `domain` but a `name`, and the record's `name` is less than 0.9 similar to it (see [Name similarity](#name-similarity)): `name_mismatch`.
4. Any of `name`, `domain`, `employee_range` or `industry` is missing or empty, checked in that order: `field_missing:<field>`, for example `field_missing:industry`.
5. Otherwise: pass. `country` is optional and is not checked. When you sent a domain, the name is not compared.

### enrich_person

**Passes if:** Name (fuzzy ≥ 0.9) and company domain (exact) match the input; a title and one contact field are present.

enrich_person input

| Input field | Meaning |
|---|---|
| first_name | string |
| last_name | string |
| company_domain | company domain |

enrich_person output

| Output field | Meaning |
|---|---|
| name | string |
| company_domain | string |
| title | string |
| email | optional |
| phone | optional |
| linkedin_url | optional |

1. No result object, or an empty one: `no_result`.
2. The record's `name` is less than 0.9 similar to `first_name` + space + `last_name` from your input (see [Name similarity](#name-similarity)): `name_mismatch`.
3. The record's `company_domain` is not your `company_domain` after normalisation: `company_mismatch`.
4. `title` is missing or empty: `field_missing:title`.
5. None of `email`, `phone` or `linkedin_url` is present: `field_missing:contact`. A person you cannot contact is not the result you paid for.
6. Otherwise: pass. One contact field is enough; the rule does not verify the email or phone it contains.

### web_search

**Passes if:** At least N results (5 by default) with valid, non-duplicate URLs.

web_search input

| Input field | Meaning |
|---|---|
| query | string, up to 500 characters |
| n | 1–25, default 5 |

web_search output

| Output field | Meaning |
|---|---|
| results | [{ url, title, snippet? }] |

1. `n` is your input's `n`, or 5 when you did not send one.
2. Every item in `results` with a `url` that parses as an `http` or `https` URL is counted once. Two items with the same URL count as one; an item with no URL or an unparseable one is not counted.
3. No usable URL at all: `no_result`, a fail.
4. At least `n` usable URLs: pass.
5. Some but fewer than `n`: partial, with the reason `results:<found>/<n>` and the fraction `found ÷ n`. See [Partial results](#partial-results).

`title` and `snippet` are not checked. A result with a valid URL and no title still counts.

### extract_url

**Passes if:** HTTP 200 with at least 200 characters of main content, not a block or captcha page.

extract_url input

| Input field | Meaning |
|---|---|
| url | http(s) URL |

extract_url output

| Output field | Meaning |
|---|---|
| url | string |
| status_code | number |
| title | optional |
| content | main content as text/markdown |

1. No result object: `no_result`.
2. `status_code` is not 200: `http:<status_code>`, for example `http:403` or `http:404`; `http:none` when the provider gave no status at all.
3. The first 2,000 characters of `content` contain `captcha`, `access denied`, `are you a robot`, `enable javascript`, `cloudflare` or `request blocked` (case does not matter) and the whole content is under 5,000 characters: `blocked_page`. A long article that merely mentions one of those words is not a block page.
4. `content` has fewer than 200 characters: `content_too_short`.
5. Otherwise: pass. `title` is optional and is not checked.

### Name similarity

`enrich_company` (when you sent a name and no domain) and `enrich_person` compare names with a similarity score from 0 to 1 and require at least 0.9. Both names are lower-cased and everything that is not a letter `a` to `z` is removed, so spaces, dots, hyphens and case never matter: `Emily Carter` and `emily-carter` are identical. The score is 1 minus the Levenshtein edit distance divided by the length of the longer string. `Emily Carter` against `emily.carter` is 1.0 and passes. `Jason Miller` (11 letters) against `Jason Millerr` (12 letters) is 1 − 1 ÷ 12 ≈ 0.92 and passes. `Amy Ross` (7 letters) against `Amy Rose` (7 letters) is 1 − 1 ÷ 7 ≈ 0.86 and fails with `name_mismatch`. Short names are therefore strict: a one-letter difference in a seven-letter name is a mismatch.

## Partial results

Only `web_search` can be partial. You ask for `n` results (the default is 5). If the provider returns at least `n` valid, distinct URLs, the item passes and the full price is charged. If it returns some but fewer, you get all of them and pay `price × found ÷ n`, rounded up to the next whole credit; the rest of the hold is released. The answer's `status` is `partial` and its `reason` is `results:<found>/<n>`. No usable URL at all is a fail with `no_result`, not a partial.

Pro rata charges on a 10-credit web search with n = 5

| Found | Fraction | Charged | Refunded | status |
|---|---|---|---|---|
| 5 or more | 1 | 10 | 0 | `passed` |
| 4 | 0.8 | 8 | 2 | `partial` |
| 2 | 0.4 | 4 | 6 | `partial` |
| 1 | 0.2 | 2 | 8 | `partial` |
| 0 | 0 | 0 | 10 | `failed` (`no_result`) |

Rounding is up, so 2 of 5 on a 7-credit tool is `ceil(7 × 0.4)` = 3 credits, not 2.8. With a test key the same fraction is reported as `would_have_charged`.

**Response (example): 2 of 5 results on a 10-credit tool**

```json
{
  "status": "partial",
  "execution_id": "5e8c1a7b-2d3f-4a6e-8b9c-1f2e3d4c5b6a",
  "receipt_id": "a1b2c3d4-e5f6-4a7b-8c9d-0e1f2a3b4c5d",
  "tool_id": "exa-web-search",
  "task_type": "web_search",
  "check_version": "v1",
  "charged": { "credits": "4", "usd": "0.004" },
  "refunded": { "credits": "6", "usd": "0.006" },
  "result": {
    "results": [
      { "url": "https://example.com/saas-directory", "title": "US SaaS companies", "snippet": "A directory of…" },
      { "url": "https://acme.com/blog/saas-in-the-us", "title": "SaaS in the US", "snippet": "The market…" }
    ]
  },
  "reason": "results:2/5"
}
```

## Our own verifier for found emails

A finder that reports its own address as valid is grading its own work. So for `find_email`, with a live key, Arettic does not take the finder's word for it. Once the finder returns an address, Arettic sends that address to its own verifier, a `verify_email` tool from the catalog that is not the tool you bought from, and writes the verifier's `status` into the result's `verification_status`. The `find_email` rule then judges that verdict. If the finder says `valid` and the verifier says `catch_all`, the item fails with `catch_all` and is refunded.

- The verifier's call is paid by Arettic. It is built into the tool's price and recorded on the execution as checker overhead; it is never added to your charge.
- The verifier gets 35 seconds. If it errors or times out, the status is `unknown`, the rule fails the item with `status:unknown`, and you are refunded. Arettic's verifier failing never costs you the price.
- If no verifier is configured on Arettic's side, the finder's own `verification_status` is judged instead. Either way the published rule is the same: only `valid` passes.
- With a test key there is no second call. The mock finder's own `verification_status` is judged, so a domain ending in `.catchall.test` fails with `catch_all`. See [Test mode](/docs/test-mode#make-a-mock-fail).

This is why `find_email` is the strictest rule: a passed address has been found by one provider and confirmed deliverable by another.

## Never charged: provider errors, timeouts and Arettic's own bugs

The pass rule only runs on an answer. When there is no answer, the item is released without running it, and the reason says why. None of these is ever charged, under any pricing mode.

Reasons that are released before or instead of the check

| reason | When |
|---|---|
| provider_error | The provider failed on both attempts: a 5xx, a 429, another 4xx such as a rejected key, a connection failure, or no adapter configured for the tool. A provider that answers 404 or 422 is not an error: that is an empty answer, and the rule fails it with `no_result`. |
| timeout | No answer within 30 seconds, on both attempts. |
| check_error | The rule itself threw on this result. Counted as a fail so you never pay for Arettic's bug. The check is recorded on the execution so it can be fixed. |
| interrupted | The process died while the item was mid-call. A safety net releases the hold within a few minutes. |
| cancelled | The item's job was cancelled, or the hold was force-settled, before it ran. |

A provider error or a timeout also counts against the tool's reliability score, and a run of them pauses the tool, so a failing provider is routed around rather than paid. If you sent `fallback: true`, a failed check, a provider error or a timeout is retried once on the next-ranked tool within your `max_price`; see [Execute](/docs/execute#fallback).

## check_version: how rules change

Every rule has a version, currently `v1` for every rule. The version that judged a result is the `check_version` on the execute answer and on the receipt, at the top for the request and on every item in `items[]`. A receipt keeps the version that was applied at the time, for ever: if the rule for a task type moves to `v2` later, a receipt checked under `v1` still says `v1`, and a dispute on it is decided against `v1`.

- A rule changes only by a new version. The text and the code of `v1` do not change once published.
- A new version is announced on the [changelog](/changelog) under the Pass rules area, with what changed and from when, before it takes effect. The changelog is also available as Markdown, JSON and RSS.
- From that date, new executions of that task type carry the new version. Executions and receipts before it are untouched.
- The `pass_rule` on every tool (`task_type@version`) and on `GET /v1/task-types` moves with it, so an agent that reads the version from the API sees the change without reading the changelog.

There is no way to ask for an older version on a new request. The current rule is the only rule a new result is checked against.

## Read the rules from the API

The rules are open data (CC BY 4.0) and need no key. `GET https://api.arettic.com/v1/task-types` returns one entry per task type with its `pass_rule` name, `passes_if`, `input` and `output`. `GET https://api.arettic.com/v1/formula` returns the score formula and its rules, and under `pass_rules` the same rules with `partial_or_failed` as well. Both carry the `license` and a `generated_at` time, and answers are cached for 60 seconds. [Public data](/docs/public-data) has every open endpoint.

**curl**

```bash
curl https://api.arettic.com/v1/task-types

curl https://api.arettic.com/v1/formula
```

**TypeScript**

```ts
import { Arettic } from "@arettic/sdk";

const arettic = new Arettic(); // no key needed for public data

const { task_types } = await arettic.taskTypes();
for (const t of task_types) console.log(t.pass_rule, t.passes_if);

const formula = await arettic.formula();
console.log(formula.pass_rules); // every rule with partial_or_failed
```

**Python**

```python
from arettic import Arettic

client = Arettic()  # no key needed for public data

for t in client.task_types()["task_types"]:
    print(t["pass_rule"], t["passes_if"])

print(client.formula()["pass_rules"])  # every rule with partial_or_failed
```

Over MCP the `list_task_types` tool takes no arguments and returns the same JSON as `GET /v1/task-types` as text content. Client configs are on [MCP server](/docs/mcp).

**MCP tool call**

```json
{ "name": "list_task_types", "arguments": {} }
```

**Response (example): GET /v1/task-types, one entry shown**

```json
{
  "task_types": [
    {
      "task_type": "find_email",
      "pass_rule": "find_email@v1",
      "passes_if": "An email is returned and its verification status is valid (not catch-all or unknown)",
      "input": {
        "first_name": "string",
        "last_name": "string",
        "domain": "company domain, e.g. acme.com"
      },
      "output": {
        "email": "string",
        "verification_status": "valid | invalid | catch_all | unknown",
        "confidence": "0–1, optional"
      }
    }
  ],
  "license": {
    "id": "CC-BY-4.0",
    "url": "https://creativecommons.org/licenses/by/4.0/",
    "attribution": "Arettic (arettic.com)"
  },
  "generated_at": "2026-09-29T10:00:00.000Z"
}
```

## See also

- [Execute](/docs/execute): the request, every status, and the full list of reason codes.
- [Receipts](/docs/receipts): where `check_version`, `outcome` and `reason` live on a receipt.
- [Disputes](/docs/disputes): what to do when a passed result was wrong.
- [Test mode](/docs/test-mode): the inputs that make each mock tool pass, fail or go partial.
- [Error codes](/docs/errors): `invalid_input` and every other refusal.
- [Schemas](/schemas): the JSON Schemas for execute requests and responses.

Updated 2026-09-29. This page as HTML: https://arettic.com/docs/pass-rules · Markdown: https://arettic.com/docs/pass-rules.md · JSON: https://arettic.com/docs/pass-rules.json
