Research Tool Bench

How We Test Screening Tools: Datasets, Metrics, Assumptions, and Reproducibility

Last updated 2026-08-16

This page is our pre-registration. It is published before we run anything, and it does not change after we see results.

If you only read one sentence, read this one: we tell you exactly what we did not measure. Most “best screening software” pages compare marketing pages. We would rather publish a small table with an honest boundary around it than a large table that quietly guesses.


Why this page exists

Choosing screening software is a decision with a real cost. Pick a tool that misses relevant records and you damage the review. Pick a tool your budget cannot renew and you lose your project mid-screening.

The information you would need to make that decision well is scattered: vendor pages describe features but not performance, academic papers report performance but on datasets and metrics that differ from each other, and pricing changes without announcement.

We try to hold those three things in one place — and, critically, to keep them separated rather than blended into a single misleading score.

Our commitment on independence

Benchmark results and recommendations are independent of monetization. A free or non-affiliate product must rank first whenever the prespecified evidence shows it performs best.

Concretely:

The three layers, and why we never merge them

Every claim on this site belongs to exactly one of three layers. Blending them is the single most common way tool comparisons become fiction.

Layer A — Directly tested

Tools we could run ourselves, on the same dataset, under the same protocol, with the same metrics. Only these appear in the quantitative leaderboard.

Layer B — Published evidence

Tools we could not run ourselves. We summarise and re-analyse verifiable published studies, recording for each: the tools and versions evaluated, dataset, the study’s own metric definitions, sample, whether the vendor was involved in the evaluation, and stated limitations.

These results are reported separately. They are never placed in the Layer A leaderboard, because a number produced under a different protocol is not comparable, however similar its name.

Layer C — Feature and price verified

Pricing, plan limits, and trial restrictions, taken only from the vendor’s own pages, with the date we checked. Third-party review sites are not used as a source for prices.

The empty cell rule

If a tool is not in Layer A, its performance cell stays empty and is marked NOT DIRECTLY TESTED, with the reason.

We will not estimate. For example, Covidence’s individual trial is limited in the number of records it accepts, which makes a full-corpus run impossible for us. The honest output of that situation is an empty cell and an explanation — not an inferred number.

Dataset

Layer A runs use SYNERGY (doi:10.34894/HE6NAQ), an open dataset of study selection in systematic reviews: 26 reviews, 169,288 records, 2,834 of them labelled as included (1.67%). It is public, citable, and re-runnable by anyone who disagrees with us.

Subset selection. We do not run all 26 reviews and then choose which to report. We specify the subset in advance, and we report every review we specified — including the ones where results are unflattering or ambiguous. Selection criteria, fixed before running:

The exact review list, with the reason each was chosen, is published alongside the results.

Metrics

Metric Definition
Recall @ 5% / 10% / 20% screened Share of all relevant records found after screening that fraction of the corpus
Records to 95% recall How many records you must read to find 95% of the relevant ones
Records to 100% recall Same, for every relevant record
WSS @ 95% Work Saved over Sampling: 0.95 − (records to 95% recall ÷ total records)
Setup cost Wall-clock time and steps required to get from zero to a running screen

Why not recall at 50%? We dropped it. Simulations stop once the last relevant record is found, so recall at high screening fractions is almost always 1.0 and separates nothing. It looks like a measurement and carries no information.

Why “records to recall” matters more than a percentage. “WSS 0.83” is abstract. “You read 247 records instead of 2,000” is the number that changes whether you can finish the review this month.

Runs, seeds, and variance

We never publish a single run.

In our pipeline test, changing only the random seed moved recall at 10% screened from 0.767 to 0.833 on identical data with an identical tool — a difference of 8.6 percentage points produced by nothing but chance. A single-run leaderboard would be reporting luck as if it were performance.

Therefore, for every tool × review combination:

Where the range for two tools overlaps substantially, we say they are not distinguishable on this evidence. We do not rank inside noise.

Stopping rules, exclusions, missing data

Stopping. Simulation runs until the last relevant record is found, or until the corpus is exhausted. Records after the stopping point count as unscreened.

Tool inclusion. A tool enters Layer A only if it can be run non-interactively on our corpus with a documented, reproducible configuration. Tools requiring manual clicking through a web interface for tens of thousands of records are not excluded from the site — they are documented in Layers B and C.

Exclusions we declare in advance. A run is excluded only for a mechanical failure — crash, timeout, corrupted output — and every exclusion is listed with its cause and the failing seed. Results are never excluded for being unexpected.

Missing data. Records lacking a title or an abstract are retained and counted, because they exist in real reviews and how a tool copes with them is part of its performance. The proportion of such records per corpus is reported.

Reproducibility

Everything needed to repeat our work is public: run scripts, configuration, tool and model versions, seeds, and raw results as CSV. The command to reproduce a result sits at the top of the results page.

If we cannot reproduce a result ourselves, we do not publish it.

Conflicts of interest

This site is run by a practising clinician-researcher who has conducted systematic reviews and has a direct interest in these tools working well. That is the source of whatever judgement this site offers, and it is also a bias worth naming.

The site intends to earn revenue through affiliate links to some of the tools discussed. The rules above exist to keep that revenue from touching the results. Every commercial link is disclosed. Where a top-ranked tool pays us nothing, we say so on the page.

What would change our conclusions

We would revise a published result if: a tool releases a version that materially changes its screening behaviour; someone reproduces our runs and gets different numbers; we discover an error in our own configuration; or a vendor demonstrates that our setup misused their tool.

Corrections are published as corrections, dated, with the previous value visible. We do not silently edit results.

Limitations we already know about