How scores work
The score formula, its six inputs, where each comes from, when scores change and how to cite one.
Every tool gets a score from 0 to 100 for each task type and region it serves. It says how likely the tool is to give your agent a correct, usable result. Scores come from our own benchmarks against test sets with known answers, plus what happens on live purchases. Nobody can pay for a better score: whether a tool can be bought through Arettic, or its provider pays us for analytics, never changes its rank.
The formula (v1)
Score = 100 × (0.40A + 0.25S + 0.20P + 0.10R + 0.05L) − min(20, 500D), clamped to 0–100
| Input | Weight | What it is | Where it comes from |
|---|---|---|---|
| A, accuracy | 0.40 | Share of benchmark cases where a correct result was delivered. | The tool's latest benchmark run: each case's result is compared with the test set's known answer. |
| S, pass rate | 0.25 | Share of calls whose result passed the pass rule. | Benchmark cases plus live calls in the last 28 days, weighted by volume. |
| P, audit precision | 0.20 | Share of audited passes a reviewer confirmed correct. | The weekly audit. Until a tool has 50 audited results, P uses A. |
| R, reliability | 0.10 | 1 − the share of provider errors and timeouts. | Benchmark and live calls in the last 28 days. |
| L, speed | 0.05 | The task type's median latency ÷ this tool's latency, capped at 1. | Benchmark p50 latency, compared with the other tools for the same task type and region. |
| D, disputes | penalty | Upheld disputes ÷ passed live results. | Live purchases in the last 28 days. Each 0.1% of upheld disputes costs half a point, up to 20 points. |
Rules
- Small samples are cautious. A rate from fewer than 200 data points uses the 95% Wilson lower bound, so a tool with 9 passes out of 10 doesn't outrank one with 900 out of 1,000.
- Scores move slowly. A score moves at most ±10 points a week, unless the tool is paused (then it can drop freely).
- Weekly. Scores are recomputed every Monday at 00:00 UTC. Benchmarks re-run on the 1st of every month.
- Per task type and region. A tool that verifies emails well in the US may not in the UK; each gets its own score, and each score says how many data points (
sample_size) it rests on. - Ranking.
recommendranks by score; within 2 points, the cheaper tool goes first. Every score input is in the answer. - A floor for sale. A tool can only be bought through Arettic after a recent benchmark with accuracy of at least 60%.
Get scores
GET https://api.arettic.com/v1/scores: the current score of every tool (filter withtask_typeandregion).GET https://api.arettic.com/v1/tools/{id}: one tool with its score per region, every input, the weekly history and its latest benchmark.GET https://api.arettic.com/v1/formula: the formula, weights, inputs, rules and pass rules as JSON.- The scores page and each tool's page show the same, and so do their
.mdand.jsoncopies. No key needed; see public data.
curl "https://api.arettic.com/v1/scores?task_type=verify_email®ion=US"
Citing a score
Scores are published under CC-BY-4.0: use them anywhere, with the attribution "Arettic (arettic.com)". Cite the tool, the task type and region, the score, the week it was computed, the formula version and the sample size, for example: "ZeroBounce, verify_email (US): 87.4, week of 2026-09-28, formula v1, n = 1,240. Source: Arettic (arettic.com)."
Check it yourself
The formula has an independent reference implementation in Python (harness/arettic_harness/formula.py, published with the benchmark harness), tested against the same cases as ours. With the published inputs of any score, you can recompute it.
Scores in test mode
With a test key, recommend returns only the mock tools, one per task type. Mock tools never appear in public scores (unless you ask /v1/tools for mode=test). See test mode.
Private benchmarks
On the Pro plan you can benchmark tools on your own test set: upload cases with known answers (POST /v1/orgs/{orgId}/test-sets), pick tools, and run (POST /v1/orgs/{orgId}/benchmarks). The results use the same pass rules and accuracy checks as ours and are private to your org: they never change public scores. 500 calls a month are included; beyond that, calls are charged at the provider's list price × your plan's multiplier, pass or fail.