Buying results

How scores work

The score formula, its six inputs, where each comes from, when scores change and how to cite one.

Every tool gets a score from 0 to 100 for each task type and region it serves. It says how likely the tool is to give your agent a correct, usable result. Scores come from our own benchmarks against test sets with known answers, plus what happens on live purchases. Nobody can pay for a better score: whether a tool can be bought through Arettic, or its provider pays us for analytics, never changes its rank.

The formula (v1)

Score
Score = 100 × (0.40A + 0.25S + 0.20P + 0.10R + 0.05L) − min(20, 500D), clamped to 0–100
The six inputs, each a rate from 0 to 1
InputWeightWhat it isWhere it comes from
A, accuracy0.40Share of benchmark cases where a correct result was delivered.The tool's latest benchmark run: each case's result is compared with the test set's known answer.
S, pass rate0.25Share of calls whose result passed the pass rule.Benchmark cases plus live calls in the last 28 days, weighted by volume.
P, audit precision0.20Share of audited passes a reviewer confirmed correct.The weekly audit. Until a tool has 50 audited results, P uses A.
R, reliability0.101 − the share of provider errors and timeouts.Benchmark and live calls in the last 28 days.
L, speed0.05The task type's median latency ÷ this tool's latency, capped at 1.Benchmark p50 latency, compared with the other tools for the same task type and region.
D, disputespenaltyUpheld disputes ÷ passed live results.Live purchases in the last 28 days. Each 0.1% of upheld disputes costs half a point, up to 20 points.

Rules

Get scores

curl
curl "https://api.arettic.com/v1/scores?task_type=verify_email&region=US"

Citing a score

Scores are published under CC-BY-4.0: use them anywhere, with the attribution "Arettic (arettic.com)". Cite the tool, the task type and region, the score, the week it was computed, the formula version and the sample size, for example: "ZeroBounce, verify_email (US): 87.4, week of 2026-09-28, formula v1, n = 1,240. Source: Arettic (arettic.com)."

Check it yourself

The formula has an independent reference implementation in Python (harness/arettic_harness/formula.py, published with the benchmark harness), tested against the same cases as ours. With the published inputs of any score, you can recompute it.

Scores in test mode

With a test key, recommend returns only the mock tools, one per task type. Mock tools never appear in public scores (unless you ask /v1/tools for mode=test). See test mode.

Private benchmarks

On the Pro plan you can benchmark tools on your own test set: upload cases with known answers (POST /v1/orgs/{orgId}/test-sets), pick tools, and run (POST /v1/orgs/{orgId}/benchmarks). The results use the same pass rules and accuracy checks as ours and are private to your org: they never change public scores. 500 calls a month are included; beyond that, calls are charged at the provider's list price × your plan's multiplier, pass or fail.

Updated 2026-09-30 · This page as Markdown · JSON