How to Test an Email Verification Service Yourself
Build a control set of addresses whose real status you already know, run it through any vendor, and measure. Here is the method, including the traps that make most homemade accuracy tests meaningless.
Every verification vendor publishes an accuracy figure. We say 98%. None of us, ourselves included, can hand you an independent third-party audit of that number. The figures are self-reported, the test sets are private, and the methodologies are not published.
You do not have to accept that. You can build a small control set where you already know the correct answer for every address, run it through any vendor, and measure. It takes an afternoon and it produces a number that describes your data rather than someone's marketing page.
Here is the method, and the mistakes that make most homemade tests worthless.
01Why you need a control set rather than a sample
The instinct is to take 1,000 addresses from your list, run them, and see what comes back. That tells you the distribution of verdicts. It cannot tell you whether the verdicts are correct, because you have no independent truth to compare against.
A control set is different: it is addresses whose real status you already know, from evidence outside the verification tool. Then every verdict is either right or wrong, and you can count.
The whole difficulty of this exercise is assembling that known truth honestly.
02Building the control set
Aim for 200 to 400 addresses across five categories. You want enough of each to be meaningful and few enough that you can source them properly.
Known good (target 80 to 150)
Addresses you have proof are live. The proof standard matters: not "we mailed it once and it did not bounce," but a delivered message with a reply, a click, or a confirmation event in the last 30 days.
Your own team's addresses count. Recent customer replies count. Anyone who clicked a link last week counts. Use a mix of providers: Gmail, Microsoft 365, Yahoo, and at least a few small business domains, because provider behaviour differs enormously and a control set that is 90% Gmail tests one code path.
Known bad (target 50 to 100)
Addresses you have proof are dead, from hard bounces in the last 30 days with the SMTP response recorded. A 550 5.1.1 on a recent send is solid evidence the mailbox does not exist.
Recency matters more here than anywhere. A bounce from eight months ago is not evidence about today.
Deliberate invalid (target 20 to 30)
Ones you construct: random strings at real domains, like [email protected]. Nobody owns that mailbox. Any service marking it deliverable is guessing.
This category is genuinely useful because you have certainty without needing a bounce record.
Known catch-all (target 20 to 30)
Domains configured to accept mail to any address. If you know a domain is catch-all, put both a real address and a fabricated one at that domain into the set.
This category is the most revealing part of the whole test, for reasons I will come to.
Known disposable (target 20 to 30)
Addresses at temporary-mail domains. Easy to generate: go to a throwaway-mail service and take the domain.
03Running the test
Three rules, and skipping any of them invalidates the result.
Shuffle the file and strip anything identifying. No column saying expected_status. Randomise the row order. If your control set arrives with 100 known-bad addresses grouped together, you have created a pattern.
Run every vendor within the same short window. Mailbox state changes. If you test one vendor in March and another in June, the difference you measure includes three months of decay. Same day is ideal.
Test the tier you would actually buy. Several vendors run a shallower pipeline on free tiers, so a free-tier test tells you about a product you are not purchasing. Ask directly whether the free tier includes mailbox-level checks. If it does not, pay for the smallest plan and test that instead.
04Scoring, and the part everyone gets wrong
Now count. And here is where most homemade tests go off the rails.
The naive score is: correct verdicts divided by total addresses. That is wrong, and it is wrong in a way that systematically rewards the worst behaviour.
Consider your known catch-all domains. You have a real address and a fabricated one at the same catch-all domain. The truthful verdict for the fabricated one is risky or unknown, because the server accepts everything and acceptance carries no information. A vendor that says "deliverable" is guessing, and it will be right about half the time by luck.
If you score "risky" as a miss, you have penalised the vendor for telling you the truth and rewarded the one that guessed. Your test now selects for overconfidence, which is the exact failure mode you were trying to detect.
So score in three buckets:
- Correct: the verdict matches known reality
- Wrong: the verdict contradicts known reality, meaning deliverable on a known-dead address or undeliverable on a known-live one
- Declined: the vendor reported risky or unknown
Then compute two separate numbers:
Accuracy on decided addresses = correct / (correct + wrong)
Decision rate = (correct + wrong) / total
These trade off against each other, and that is the actual finding. A vendor can hit near-perfect accuracy by declining anything uncertain. Another can decide everything and be wrong more often. Neither is straightforwardly better, and which you want depends on what you do with the results.
The single most important number in the whole exercise is the wrong count on known-dead addresses. Every one of those is a hard bounce you will generate in production. That is the failure with a real cost.
05What to look for in the results
False deliverable on known-dead addresses. The expensive error. Each one is a bounce against your sender reputation.
False undeliverable on known-good addresses. The quieter error, and it is not free. Each one is a real person you delete from your list. If a vendor marks your own colleague's working address as undeliverable, ask why before trusting the rest.
Behaviour on the deliberate invalid addresses. A random string at gmail.com should never come back deliverable. If it does, the service is not performing a mailbox-level check on that provider.
Behaviour on catch-all domains. Does the fabricated address at a catch-all domain get flagged as uncertain, or confidently marked valid? This tells you more about a vendor's honesty than any published figure.
Disposable detection rate. Straightforward hit rate, and expect variation, because the domain lists change constantly.
06Honest limits of this method
Three, and they matter.
200 to 400 addresses is a small sample. At that size your confidence interval is wide, several percentage points at least. This test is good at spotting large differences between vendors and bad at resolving small ones. If two services come out within three points of each other, you have learned they are comparable, not which is better.
Your control set reflects your data. If your list is 95% Gmail, you have measured Gmail performance. That is arguably what you want, since it is your list, but do not generalise it.
Known-bad addresses go stale. Mailboxes get reactivated. A bounce from last week is strong evidence; from last year it is not evidence at all.
07What this does not measure
Speed at volume, which you should test separately with a large real file. CSV handling and whether your other columns survive. What happens to your data afterwards, which is a contractual question. And whether uncertain verdicts are actionable in your workflow.
Our post on choosing a verification service covers those criteria, which are not accuracy questions and often matter more.
08Why we are telling you how to do this
Because we claim 98% and cannot independently prove it, and we would rather you tested it than took our word.
The result we expect on a well-built control set is a high accuracy figure on decided addresses and a visible decline rate, because we report risky and unknown rather than guessing on catch-all domains and unresponsive servers. That will look worse than a competitor that decides everything, on a naive score. It should look better on the wrong-count, which is the number that costs you money.
If your test says otherwise, we want to know. That is a more useful thing to have than an unfalsifiable percentage.
Run your control set. 500 free credits a month on the full deep scan, which is more than a 400-address control set needs.
Start with 500 free validation credits. No card.
Both Free and Pro run the same scan engine, full SMTP probe, MX lookup, typo, disposable, domain checks, and the evidence chain on every verdict. The difference is the monthly credit pool (Free=500, Pro=10,000, Max=75,000) plus Pro's API and MCP access.