Most advice on vetting prospects is a list of criteria. Criteria are the easy part. The expensive part is the order you apply them in, because every check costs time and most of them return a no. Get the order wrong and you spend an afternoon reading pages you could have rejected in a second.
This is the order that survives contact with a real list. It comes out of opening all 1,001 rows of a purchased list by hand and recording what each rejection actually cost: the full run is published, with denominators.
The rule underneath all of it: reject on the cheapest signal that can reject, and never spend a page-load on something a string match could have caught.
Stage 1: reject without loading anything
These cost nothing but a pass over the file, and both of them fired hard on the list we measured. How much they remove together depends on how far the two overlap on your file, which is why the figures below are given separately rather than added up.
Deduplicate by host, not by URL
Our list had 1,001 rows across 467 hosts, and 69.9% of rows sat on a host that appeared more than once. One Polish forum contributed 30 rows on its own. A list is always smaller than its row count, and if you are going to write to a human, twelve pages on one domain is one conversation, not twelve.
Group by registered domain. Keep the best two or three rows per host and park the rest. This is the single highest-yield operation in the whole process and it takes one spreadsheet formula.
Reject page types that never take a link
Some URLs are recognisable as dead ends from the path alone. On our list, audio and video players, forum threads and guestbooks were 243 rows, 24.3% of the list, and almost none of them could carry a comment.
Patterns worth rejecting before you fetch anything:
/tag/,/category/,/author/and other archive paths. An archive is a list of other pages./sermons/,/media/,/podcast/,/watch/and player paths generally./forum/,/topic/,viewtopic.php. A forum thread is not a blog post and a reply is not a comment./guestbook/,/gastebuch/. Still surprisingly common on bought lists.- Anything with a session id or a tracking parameter in it. That URL will not be stable.
Reject the obviously wrong language and country
If your campaign is English and the URL is on a .ru or .pl domain with a Cyrillic or Polish slug, that is a decision you can make now for free. Do not defer it to the reading stage.
Stage 2: check in bulk, before a human reads anything
Everything in this stage is mechanical. It is the part worth automating, and it is what the Backlink Prospect Checker exists to do, but the order matters whether you use a tool or write your own script.
Does the page load at all
Start here because it is the cheapest request and it invalidates everything downstream. On our list 213 of 1,001 rows, 21.3%, could not be read: dead, blocked, password-protected, or in a redirect loop. A fifth of a purchased list is not a page.
Keep the failures in your denominator. A row that could not be read is not a row you get to ignore when you report on list quality, and quietly dropping them is how “we found a 40% success rate” gets published.
Is the placement you want actually possible on this page
For comment links, that means: is there a comment form, on this exact page, today. It is a page-level and template-level property, not a domain one. Comments get closed when a post is archived, a theme change removes the form, and a plugin can turn the whole thing off without editing a word of the copy.
Whether a comment link is worth pursuing in the first place is the prior question, and worth settling before you spend a stage on it. This stage only establishes whether the page could take one.
On our list, 183 of 1,001 rows, 18.3%, had an open form. 818 had none at all.
What the placement would be worth
An open form is not a followed link. Read what the page already gives its commenters rather than assuming a platform default. Of the 183 open forms, only 37 would have carried a followed link: 91 were nofollow, 22 were ugc plus nofollow, and 33 had no existing comment links to read at all and so could not be judged either way.
That last group matters. An open form with no comments is genuinely unknown, and a tool that reports it as “dofollow, probably” has put an assumption in a column that looks like an observation. See the guide on link attributes for how to read it properly.
What is already sitting in the comments
The existing thread is the cheapest quality signal on the page. Fifty comments of pharmacy links means the moderation is not moderating, and it tells you what your link would be sitting next to. On our list 50 rows, 5.0%, came back High spam risk and 132, 13.2%, Medium.
Stage 3: the three judgements a person still has to make
Everything above can be automated. These cannot, and pretending otherwise is where prospecting tools lose credibility.
Is this page relevant to what you are building links for
Relevance is a property of a pair: this page and your site. No tool can score it from the URL alone, because the same page is an excellent prospect for one campaign and noise for the next.
Practical version: write your subject in five specific words, then sort into three buckets, same subject, adjacent, unrelated. Three buckets can be applied consistently across a thousand rows. A ten-point relevance score cannot, and it drifts as you get tired.
Is this a site you want your name next to
Not a metric. Look at the last three posts, not the tagline. The tagline says what the site meant to be; the recent posts say whether anyone is still there. On our list, 163 rows, 16.3%, were unedited theme-demo posts, the sample articles that ship with an e-commerce template and never get deleted. They are technically live pages with technically open comments and no human has ever read one.
Is the effort proportionate to what you would get
A moderated comment on a real site with a real audience is worth writing carefully. A followed link on a theme-demo post at DR 2 is worth roughly nothing, and it will still take you fifteen minutes. On the list we measured, exactly ten rows out of 1,001 were followed, readable and low spam risk at the same time. That is the number that decides whether the list was worth buying.
The order, as a table
| Stage | Check | Cost | What it removed on our list |
|---|---|---|---|
| 1 | Group by host | None | 70% of rows were duplicates at host level |
| 1 | Reject dead page types by URL pattern | None | 243 rows, 24.3% |
| 2 | Can the page be read | One request | 213 rows, 21.3% |
| 2 | Is the placement possible | Same request | 818 rows had no comment form |
| 2 | What would the link carry | Same request | 113 of the 183 open forms passed nothing |
| 2 | Spam risk in the existing thread | Same request | 182 rows, 18.2%, Medium or High |
| 3 | Relevance to your campaign | A person | Cannot be automated |
| 3 | Is this a real site | A person | 163 rows were theme demos |
| 3 | Is the effort proportionate | A person | 10 rows survived everything |
All figures from one purchased list of 1,001 rows, 999 distinct URLs across 467 hosts, opened and verified by hand on 7 September 2026. One list is one list: treat these as this list’s rates, not as benchmarks for yours.
What this costs you by hand
Two minutes a page is a fair average once you include the pages that hang, the ones that redirect three times and the ones you have to view source on. At that rate, stages 1 and 2 across 1,001 rows is about 33 hours, to arrive at ten rows worth an email.
That arithmetic is the entire argument for automating stage 2. It is not that a machine judges better than you do. It is that you should not be spending 33 hours to find out that 818 pages have no comment form.
Where the tool fits
The Backlink Prospect Checker does stage 2 and nothing else. You give it the list, it opens every exact page and reports comment opportunity, likely link type with the evidence behind it, Ahrefs DR and comment spam risk. Stage 1 is a spreadsheet job and stage 3 is yours.
It will not find prospects for you, it does not post anything, and where it has no evidence it reports Unknown instead of a platform default. Every field’s method is written out, including what it cannot prove.
If you want to see what the output looks like before signing in, there is a full sample report.
This framework runs differently depending on where the list came from. There is a start-to-finish version for an Ahrefs export you pulled yourself, where stage 1 is mostly deduplication, and one for a list you bought, where the first thing to establish is whether the seller opened any of the pages.