>>> Blog
Inclusion and Exclusion Criteria: What to Include and What to Leave Out
Your protocol is due, and the line your supervisor circled says inclusion and exclusion criteria. Every record you screen from here on will be judged against it, and rewriting it halfway means explaining the change in your paper.
Here is the short answer. In a systematic review, inclusion and exclusion criteria are one set of rules about studies: who or what was studied, what was examined, how the study was done, and which kinds of report count. Two more appear only when the question needs them: a comparison, and outcomes. Write the rules before you screen, in words a second reviewer could apply without asking what you meant.
The map below shows the four criteria most reviews need, for an invented review: how do employees experience the AI tools their organisation uses at work? The columns give the criterion, what you define for it, one example of what gets in and one example of what stays out.
Examples of inclusion and exclusion criteria: one review, four criteria
| Criterion | What you define | Include | Leave out |
|---|---|---|---|
| Who or what was studied | Who counts, and where | Employees in organisations that use AI tools | Students, self-employed people |
| What was examined | The topic, in specific terms | How employees experience AI tools in daily tasks | Technical tests of how accurate a tool is |
| Study design or method | Features, not labels | Interviews, surveys, case studies | Opinion pieces |
| Kind of report | Type, language, years | Peer-reviewed papers, conference proceedings, English | Editorials, news items |
You may have met acronyms for organising these, such as PICO or SPIDER. PRISMA 2020 notes that PICO is commonly used for reviews of interventions, and what it asks you to report is every characteristic that decided eligibility (PRISMA 2020 explanation, item 5). Use a frame if it fits your question. The sections below do not depend on one.
1.Who or what was studied
People, organisations, cases, documents: whatever your studies are about, say exactly which ones count.
| Define | Who or what counts, how you tell (age, role, sector, size), and the setting |
| For example | Secondary school teachers · warehouse operators · suppliers with fewer than 50 staff · open-source projects with several active contributors |
| Watch for | Studies that only partly overlap: a study of suppliers with 20 to 80 staff only partly fits "fewer than 50", so settle the rule now |
Avoid a criterion that shuts out studies for no reason, such as a definition or measure too recent for older studies to have used.
2.What was examined
"AI at work" is a topic, not a criterion. Say what counts as an AI tool, and which experiences you mean.
| Define | The intervention, exposure, concept or experience under study, with the features that change your question: how much, for how long, delivered by whom |
| For example | A staff training programme of four or more sessions · an AI tool used to forecast demand · a named lean method on the shop floor |
| Watch for | Any limit on amount, duration or form needs a reason you can state |
3.Study design or method
Design is a trade: tighter criteria bring fewer studies at lower risk of bias, looser ones bring more at higher risk.
| Define | The kinds of study you will accept, described by what they did: how data were collected, whether there was a comparison, how long people were followed |
| For example | Surveys · case studies · field studies · interviews and focus groups · simulation studies · experiments and randomised trials · mixed methods studies |
| Watch for | Labels: the Cochrane Handbook warns that design labels are used inconsistently ("double blind" and "case-control" are its examples), so name the feature you need, not the label |
4.Kind of report
These look like housekeeping and are not: each one removes studies before anyone reads them.
| Define | Which kinds of report count (articles, theses, conference abstracts, unpublished reports), which languages you can read, and which years |
| For example | A start year that matches when the thing you review first existed, such as 2000 for a device first available then (PRISMA 2020 explanation) |
| Watch for | The Handbook wants "a compelling argument" before you exclude unpublished or grey literature, and a hard look at bias before you restrict to one language |
Write the reason next to each restriction: PRISMA 2020 suggests giving a rationale for any notable one.
5.A comparison, if your question makes one
Not every question compares. If yours does not, skip this one; SPIDER, a frame built for qualitative questions, has no letter for it.
| Define | What the thing is set against: nothing, usual practice, the previous process, another version of it, or another group |
| For example | A group with no training · the previous process · a manual method set against an AI-assisted one |
| Watch for | Vagueness: name the comparison exactly, not just "a control group" |
6.Outcomes, if your question needs them
Outcomes are the odd one out: usually they describe what you will extract, not what gets a study in.
| Define | Whether an outcome is a criterion, or only something you will extract from the studies that qualify |
| For example | Employee turnover measured, in a review of what reduces turnover · dropout measured, in a review of what keeps students enrolled · delivery delays measured, in a review of what shortens lead times |
| Watch for | A review of what prevents one particular outcome, or of an intervention's unwanted effects, is a case where the outcome does decide eligibility (Handbook, section 3.2.4.1) |
Say in the protocol which role your outcomes play, before anyone screens.
What to leave out
Some things feel like criteria and are not.
| Leave out | Why | Where it goes |
|---|---|---|
| Whether a study reports usable outcome data | It invites selective-reporting bias. Not measuring the outcome can be a reason to exclude; not reporting it should not (Handbook, section 4.6.3) | Extraction, with the gap noted |
| How good the study is | It is judged on the studies that qualify, and the protocol says how weaker ones are handled (Handbook, section 7.6.2) | Appraisal: which tool fits which design |
| A rule you thought of after reading the results | Changes need a justification and must not be driven by what the studies found (Handbook, section 3.2.1) | A documented deviation in the review |
When you record an exclusion for outcome data, say which of the two it was: PRISMA 2020 calls "no relevant outcome data" ambiguous. And a design feature you can read off the paper, such as a minimum follow-up, can be a criterion, while a judgement about how well the study was done waits for appraisal.
How to apply them
Criteria only work when two people read them the same way, so the process around them matters as much as the wording.
| Step | What to do |
|---|---|
| Write | Fix the criteria in the protocol before anyone screens |
| Pilot | Try them on six to eight reports: some clearly in, some clearly out, some doubtful. Reword until two people agree |
| Screen | Two reviewers decide each study at full text, independently, with a way to settle disagreements agreed in advance. Ideally the same at title and abstract |
| Record | Assess the criteria in order of importance, and let the first "no" be the reason for exclusion |
| Report | Methods: the criteria. Results: studies that looked eligible but were excluded, with reasons (PRISMA 2020 items 5 and 16b) |
The pilot, screening and recording steps follow the Cochrane Handbook, section 4.6.4.
In Littview, the criteria live in the project once, reviewers pick the criteria behind each decision as they screen, and the exclusion reasons a reviewer picks are what the PRISMA flow diagram counts.
FAQ
Three questions come up most.
What is the difference between inclusion and exclusion criteria? In a review there is one set of eligibility rules. Inclusion criteria say what qualifies, exclusion criteria say what disqualifies, and a study that fails a single criterion is out.
Are inclusion and exclusion criteria in clinical trials the same as in a review? Not quite, because those criteria describe participants, not studies. Inclusion criteria are the key features of the target population that the investigators use to answer their research question. Exclusion criteria in research studies describe people who meet those features but have other characteristics that could interfere with the study or raise their risk of an unfavourable outcome (Patino and Ferreira). A review turns the same idea onto studies.
Do inclusion and exclusion criteria in qualitative research work the same way? In a review of qualitative studies you still write criteria for the studies and apply them the same way. Who or what was studied, what was examined, the method and the kind of report carry over, and a comparison often has no place. SPIDER (Sample, Phenomenon of Interest, Design, Evaluation, Research type) was put forward as a frame for qualitative and mixed methods questions, though its authors say it needs refining and testing on a wider range of topics (Cooke, Smith and Booth).
For this week: take the criteria from the map, add a comparison or outcomes only if your question needs them, write each one so a second reviewer could apply it without asking you, test them on a handful of reports, and put them in the protocol before anyone screens a record.
More from the blog