Skip to content

>>> Blog

Can AI Replace Your Second Reviewer? What the Evidence Says

7 min readLittview Team

It's the question every review lead quietly asks the moment screening starts to drag: do we really need two humans on every abstract, or could AI take the second seat?

It's a fair question — dual screening is expensive, and the second reviewer's job is largely to catch a small number of errors. It's also a question with an unusually clear answer right now, because 2024–2025 produced both direct head-to-head evidence and an explicit joint position from the major evidence-synthesis organizations. Here's what they actually say.

Why the second reviewer exists at all

Dual independent screening isn't tradition for tradition's sake. Single screeners miss relevant studies — in a crowd-based randomized trial, single-reviewer abstract screening missed 13% of relevant studies on average (Gartlehner et al., 2020). Fatigue, ambiguous abstracts, and criteria drift all take a toll, and a study wrongly excluded at title/abstract stage is invisible for the rest of the review. Independent double screening with disagreement resolution is the standard guarded by Cochrane's methodological expectations (MECIR standards C39–C42) precisely because missed studies bias the evidence base in ways no later stage can repair.

So the bar for replacing the second reviewer is specific: the replacement must catch the first reviewer's misses at least as reliably as a second human would — and any studies it wrongly rejects must not silently vanish.

The direct evidence: AI in the second seat

The most on-point study to date is Tran and colleagues (2024, Annals of Internal Medicine), which evaluated GPT-3.5-class models in exactly this role — second reviewer for title and abstract screening. The results, in plain terms:

  • Under a balanced decision rule, sensitivity ranged from 81.1% to 96.5% and specificity from 25.8% to 80.4% across reviews.
  • The models produced 10,279 false positives — 45.3% of screened records — each one an abstract a human had to re-check and reject.
  • Under a sensitivity-first rule, the model missed 0–1 relevant citations — near-perfect recall, bought with even more false positives.
  • The models also caught 7 citations (about 1%) that human reviewers had missed — the second-reviewer function genuinely working.

The authors' verdict is worth quoting for its precision: LLMs "may be used as a second reviewer… at the cost of additional work to reconcile added false positives." Not no, not yesyes, with a measured price.

Newer models improve the price. A 2025 comparison found GPT-4-class specificity of 0.98 versus 0.51 for GPT-3.5 at comparable sensitivity — dramatically fewer false positives. But every one of these results was measured on particular reviews, and performance varies by topic and field. That variance is why nobody serious hands over the seat unconditionally.

The wider literature agrees on the ceiling. A 2025 scoping review in the Journal of Clinical Epidemiology covering 37 studies of LLMs in review tasks found 54% of evaluations promising — and still concluded LLMs are "not yet ready" to act as autonomous reviewers. Its search ended in February 2024, so it should be read as a broad snapshot of an early and fast-moving evidence base.

What the guidance says

In late 2025, Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence published a joint position statement — the closest thing evidence synthesis currently has to a shared policy position on this question. Three of its mandates bear directly on the second-reviewer debate:

  1. "Evidence synthesists are ultimately responsible for their evidence synthesis, including the decision to use artificial intelligence and automation." You can delegate screening labor; you cannot delegate accountability for a missed study.
  2. AI "should be used with human oversight." An AI second reviewer whose exclusions no human ever sees fails this test by construction.
  3. Any AI that makes or suggests judgments must be transparently reported — tool, version, role, and validation. The statement even provides a protocol-level reporting template, and expects you to justify the tool's use in your specific context, piloting it on your own project where certainty is low.

So: can it replace the second reviewer?

The evidence supports a nuanced answer, and the nuance is the useful part.

As a full replacement — no, not today. The direct studies, the scoping-review evidence, and the joint guidance all point in the same direction: unvalidated or fully autonomous AI screening is not ready to be treated as equivalent to a second human across review contexts, and unreported AI judgments would conflict with the joint position statement and may violate the disclosure policies of the journal involved.

As a second opinion — yes, and it's genuinely valuable. The defensible configurations today look like:

  • AI as a safety net over dual human screening — flagging papers both humans excluded that match inclusion patterns (this is where Tran's "7 missed citations caught" lives).
  • AI as a tie-breaker input — when reviewers disagree, an AI's criteria-by-criteria reasoning is a useful third signal for the human who resolves the conflict.
  • AI as the junior partner in a supervised pair — AI screens alongside one human, and a human reviews every AI exclusion (the sensitivity-first rule makes this workable: near-zero missed studies, with the false-positive reconciliation cost budgeted in).

Each configuration keeps a human answerable for every exclusion — which is the property the second reviewer was there to protect in the first place.

How Littview implements this

Littview's design matches the evidence rather than fighting it. Independent screening by multiple reviewers, with a conflict-resolution workspace for their disagreements, is the backbone of the product. Two reviewers per paper is the classic setup, but the number is yours: the project owner decides how many reviewers screen each paper, and can put three or more on contested papers where the protocol calls for extra scrutiny. However many humans are in the loop, AI doesn't replace that structure — it plugs into it:

  • The AI suggestion panel gives each screener an optional second opinion, with its reasoning and the specific criteria it invoked, before or after their own call. The decision stays yours.
  • In conflict resolution, an AI analysis of the disagreement — votes, notes, and criteria — is available to the person resolving it, as one input among the human ones.
  • Every AI suggestion is stored separately from human decisions, so your audit trail and your manuscript's AI disclosure both stay accurate.

The second reviewer's chair isn't going away. But it's getting a very capable assistant standing behind it — and, used the way the evidence supports, that assistant makes the review more rigorous, not less.


References

  • Tran V-T, Gartlehner G, Yaacoub S, Boutron I, Schwingshackl L, Stadelmaier J, et al. Sensitivity and specificity of using GPT-3.5 Turbo models for title and abstract screening in systematic reviews and meta-analyses. Annals of Internal Medicine. 2024;177(6):791–799. doi:10.7326/M23-3389
  • Oami T, Okada Y, Nakada T-A. GPT-3.5 Turbo and GPT-4 Turbo in title and abstract screening for systematic reviews. JMIR Medical Informatics. 2025;13. doi:10.2196/64682
  • Lieberum J-L, Toews M, Metzendorf M-I, Heilmeyer F, Siemens W, Haverkamp C, et al. Large language models for conducting systematic reviews: on the rise, but not yet ready for use — a scoping review. Journal of Clinical Epidemiology. 2025;181:111746. doi:10.1016/j.jclinepi.2025.111746
  • Flemyng E, Noel-Storr A, Macura B, Gartlehner G, Thomas J, Meerpohl JJ, et al. Position statement on artificial intelligence use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence 2025. Environmental Evidence. 2025;14:20. doi:10.1186/s13750-025-00374-5
  • Thomas J, Flemyng E, Noel-Storr A, Moy W, Marshall IJ, Hajji R, et al. Responsible use of AI in evidence SynthEsis (RAISE) 1: recommendations for practice. Open Science Framework. 2025. doi:10.17605/OSF.IO/FWAUD
  • Gartlehner G, Affengruber L, Titscher V, Noel-Storr A, Dooley G, Ballarini N, König F. Single-reviewer abstract screening missed 13 percent of relevant studies: a crowd-based, randomized controlled trial. Journal of Clinical Epidemiology. 2020;121:20–28. doi:10.1016/j.jclinepi.2020.01.005
  • Cochrane. Selecting studies to include in the review: standards C39–C42. Methodological Expectations of Cochrane Intervention Reviews (MECIR) Manual.