Skip to content

>>> Blog

Critical Appraisal Tools for Systematic Reviews: Which One for Which Study Design

11 min readLittview

Screening is done, the included list is final, and your supervisor asks the question nobody planned for: which appraisal tool are we using? The protocol only says that quality will be assessed.

Here is the short answer. Protocols write this step as quality assessment, risk of bias assessment or critical appraisal, and whichever words yours uses, the tool follows the design of each included study, not your preference, your field, or whatever the last review in your lab used. A randomised trial gets a trial tool, a cohort study gets a cohort tool, and a qualitative study gets a qualitative tool, even in a review that is otherwise all numbers.

This page maps each design to its tool, says what each one gives you back, and links every checklist so you can download it and start.

The map: study design to appraisal tool

Find the designs of your included studies in the first column. The second column names the tool for each, and the third says what that tool gives you back.

Study designTool(s)What it produces
Randomised trialRoB 2 · CASP RCT · EPHPPDomain judgements · Yes / No / Can't tell · Strong / moderate / weak
Non-randomised study of an interventionROBINS-I · JBI quasi-experimental · EPHPPDomain judgements, low to critical · Yes / No / Unclear · Strong / moderate / weak
Cohort or case-controlNewcastle-Ottawa Scale · CASP · JBI · EPHPPStars, up to nine · Yes / No / Can't tell · Yes / No / Unclear · Strong / moderate / weak
Cross-sectionalCASP cross-sectional · JBI analytical cross-sectionalYes / No / Can't tell · Domain-grouped Yes / No / Unclear
QualitativeCASP qualitative · JBI qualitativeYes / No / Can't tell · Yes / No / Unclear
Mixed methods, or a review mixing designsMMAT 2018Yes / No / Can't tell per criterion, no total
Diagnostic test accuracyQUADAS-3 (QUADAS-2 until 2026)Risk of bias and applicability per domain
A systematic review as the unitAMSTAR 2 · ROBISConfidence rating, high to critically low · Low / high / unclear risk of bias
Anything else: case series, case reports, prevalence, economic evaluation, prediction rulesJBI · CASP (economic evaluation, prediction rules)Yes / No / Unclear · Yes / No / Can't tell
Scoping reviewNone requiredOptional: PRISMA-ScR items 12 and 16 begin "If done"

1.RoB 2

If your review is all randomised trials, RoB 2, the Cochrane risk of bias tool for randomised trials, is the only one you need.

Use it forRandomised trials: parallel-group, cluster or crossover
You getLow risk / some concerns / high risk for each of five domains and overall, for one result at a time
Get itriskofbias.info · paper · Cochrane Handbook chapter 8

The habit to break is rating the paper. RoB 2 rates a result, so one trial can be low risk for mortality and high risk for a patient-reported outcome. Both are correct, and both go in your table.

2.ROBINS-I

The moment your included studies stop being randomised, this is where you land. It is heavier than RoB 2, and it should be: you now have to say what randomised trial the study is imitating, and where it falls short.

Use it forCohort-type, before-and-after and other studies that compare interventions without randomisation
You getLow / moderate / serious / critical risk of bias, or no information, across seven domains, judged against the randomised trial the study is trying to imitate
Get itriskofbias.info · paper · Cochrane Handbook chapter 25

"Low" here means as good as a well-run trial. Very few observational studies earn that on confounding, so do not read a run of "moderate" as a bad review. Use the 2016 tool; the V2 on the same page is still a draft.

3.Newcastle-Ottawa Scale

The star scale for cohort and case-control studies. Quick to apply, easy to read, and easy to over-interpret.

Use it forCohort and case-control studies
You getUp to nine stars: Selection (four), Comparability (two), Outcome or Exposure (three)
Get itOHRI, scale and coding manual

The scale sets no threshold. If you decide that seven stars means "good quality", that is your rule, not the scale's, and it belongs in your methods section where a reader can see it.

4.CASP critical appraisal checklists

Plain-language checklists from the Critical Appraisal Skills Programme, free to download and written for people learning appraisal, which makes them the usual starting point for anyone new to it.

Use it forRCT, cohort, case-control, cross-sectional, qualitative, diagnostic, economic evaluation, clinical prediction rule, systematic review (with or without meta-analysis)
You getYes / No / Can't tell for ten to twelve questions, with your reasons. No score
Get itcasp-uk.net

CASP for systematic reviews comes in three checklists: the general one, and two for meta-analyses of trials and of observational studies. Two habits to avoid. Do not count the Yes answers; there is no score. And do not leave the box under each question empty: a run of "Can't tell" answers is a finding about the paper, and you want it on record.

5.JBI critical appraisal checklists

The JBI appraisal tools, from the organisation that began as the Joanna Briggs Institute, are the widest set on this page: there is a checklist for almost every design, including case reports and expert opinion.

Use it forQuasi-experimental, cohort, case-control, cross-sectional, case series, case reports, qualitative, diagnostic accuracy, prevalence, economic evaluation, systematic reviews, textual evidence
You getYes / No / Unclear / Not applicable on six to eleven questions, then include, exclude or seek further information
Get itjbi.global · RCT: JBI publishes no separate checklist for it, use the paper

Say which edition you used. The quantitative tools were revised recently, older copies circulate, and a methods reviewer will notice.

6.MMAT 2018

MMAT, the Mixed Methods Appraisal Tool, is the one to reach for when your included studies are a mix of qualitative, quantitative and mixed methods work and you would rather not run three tools side by side.

Use it forQualitative, randomised, non-randomised, quantitative descriptive and mixed methods studies. Not economic or diagnostic studies
You getYes / No / Can't tell on five criteria, for the category that matches the design
Get itMMAT wiki, with the user guide

There is no overall score, and the guide is firm about that. Present the ratings per criterion and use them in a sensitivity analysis.

7.EPHPP

EPHPP is the Effective Public Health Practice Project's Quality Assessment Tool for Quantitative Studies: built for public health reviews, usable well beyond them, one tool for trials, cohort and case-control studies, with a global rating at the end.

Use it forTrials, cohort and case-control studies, before-and-after and interrupted time series designs
You getStrong / moderate / weak on six components, then a global rating: strong with no weak component, moderate with one, weak with two or more
Get itMERST, McMaster University, tool and dictionary · paper

The single global rating is what people like about it. It can appraise randomised trials alongside the other quantitative designs, which is what makes it useful for a public health review that mixes them; for a review of randomised trials only, RoB 2 gives the more specialised risk of bias assessment. It does not cover qualitative or cross-sectional designs.

8.QUADAS-3

If your review is about how well a test finds a condition, this is your tool, and it has just changed. QUADAS-3 replaced QUADAS-2 in 2026. Use QUADAS-3 for a new review, and QUADAS-2 only if your protocol already names it.

Use it forPrimary diagnostic test accuracy studies
You getRisk of bias on four domains (participants, index test, target condition, analysis) and applicability on the first three, each judged low, high or insufficient information, for one accuracy estimate at a time
Get itbristol.ac.uk · QUADAS-2 paper · QUADAS-3 paper

Both versions ask you to write down your review question and tailor the signalling questions, the yes or no prompts under each domain, before you rate a single study. Do that step first; the ratings depend on it.

9.AMSTAR 2

When the systematic review is your unit of analysis, as in an umbrella review, you appraise reviews rather than studies. AMSTAR 2 is built for reviews of interventions.

Use it forSystematic reviews of interventions, with randomised or non-randomised studies
You getYes / Partial Yes / No on sixteen items, seven of them critical, then a confidence rating: high, moderate, low or critically low
Get itamstar.ca · paper

It is not a score. One critical flaw makes the review "low", more than one "critically low", however many other items it passes.

10.ROBIS

ROBIS is the risk of bias tool for systematic reviews. It looks at the same object as AMSTAR 2 and asks a different question: not whether the review was done well, but whether its conclusions can be trusted.

Use it forSystematic reviews of interventions, diagnosis, prognosis or aetiology, in guidelines and overviews
You getLow / high / unclear concern for four domains of the review process, and an overall risk of bias
Get itbristol.ac.uk · paper

Neither AMSTAR 2 nor ROBIS is for primary studies. If you find yourself applying one to a cohort study, go back to the map.

Not appraisal tools: PRISMA, CONSORT, STROBE

A question that comes up in nearly every first meeting: can we use PRISMA to appraise the studies? No, and the reason is worth two minutes, because the same confusion catches CONSORT and STROBE too.

ConceptQuestion it answersInstruments
Risk of biasCould the result be systematically wrong?RoB 2, ROBINS-I, QUADAS, ROBIS, the revised JBI tools
Methodological qualityWas the study done to a standard?AMSTAR 2, MMAT, Newcastle-Ottawa, EPHPP, CASP, the classic JBI checklists
Reporting qualityIs the paper complete enough to judge?PRISMA 2020, CONSORT 2025, STROBE

PRISMA, CONSORT and STROBE say what a paper must contain, not whether its results can be trusted. STROBE puts it plainly: its checklist "is not an instrument to evaluate the quality of observational research". Use them when you write up.

How to apply any of them

Whichever tool you pick, the process around it is the same, and it is the process that examiners and peer reviewers check.

StepWhat to doSource
WhoTwo reviewers, independently, with a way to settle disagreements agreed in advanceCochrane Handbook 7.3.2
What to recordThe judgement for every domain or question, and the reason behind itCASP and ROBIS both give you the box for it
Where to reportMethods: the tool, how many reviewers, whether they worked independently. Results: the assessment for each included studyPRISMA 2020 items 11 and 18
What to do with itRestrict the primary analysis to low-risk studies and show sensitivity analyses, rather than dropping studies from the reviewCochrane Handbook 7.6.2

The MMAT 2018 criteria come built in to Littview's extraction template builder, one section per study design. Any other checklist goes in as custom sections and questions.

Pick the tool by design, apply it the same way to every study, with two reviewers, and say so in the paper.

FAQ

Four questions come up every time.

Which critical appraisal tool for qualitative research? CASP qualitative or JBI qualitative, ten questions each. Pick one, say which, and use it for every qualitative study in the review.

CASP or JBI? CASP asks whether you believe the results and whether they apply locally; JBI's tools focus on bias and end with an include / exclude / seek further information verdict. Reviews following JBI methodology use JBI. Either is fine if you apply it consistently and report it.

Is there a critical appraisal tool for scoping reviews? None is required. PRISMA-ScR lists appraisal as optional and JBI's guidance authors call it "not mandatory". If you appraise anyway, say why and how.

Is PRISMA a critical appraisal tool? No. PRISMA 2020 tells you what to report about your appraisal (items 11 and 18). How to appraise is the job of the tools above.

If you are starting appraisal this week: find your designs in the map, download the checklists, agree the process with your second reviewer, and write the tool names into your methods before you rate the first paper. The rest is patience.