Systematic Literature Review: What It Is and How One Is Done

  • Post category:Insights

A systematic review, or systematic literature review, is a review with three things needed before the work starts: a protocol, a reproducible search strategy, and explicit criteria for what counts as relevant evidence. The screening, the extraction, and the appraisal all elaborate those three commitments.

If that definition sounds procedural, it is. The procedure exists so that someone else, working from your published methods, should be able to arrive at the same set of included studies. A review that cannot be reproduced is a literature review with a longer bibliography.

This guide walks through what makes a review systematic, when the method is worth its cost, what the stages involve, and who does what along the way.

What makes a review systematic

Four commitments separate a systematic review from other kinds of evidence summary.

A protocol written before screening begins. The protocol states the question, the inclusion and exclusion criteria, the search strategy, and the analysis plan. Write it first and the criteria stop drifting toward the studies you happen to find. Protocols are normally registered: PROSPERO for reviews of health interventions, and the Open Science Framework or JBI for scoping reviews, which PROSPERO does not accept.

A search someone else can run. Every database, every search string, every date limit, and language restriction is recorded. Published searches are usually reproduced verbatim in an appendix. This is the part most often done badly, and the part an information specialist improves most.

Explicit criteria applied consistently. Inclusion criteria are written down before screening and applied the same way to every record. When a decision is borderline, the reasoning goes into the record with it.

Independent assessment by at least two people. Two reviewers screen the same records without seeing each other’s decisions. Disagreements are recorded and resolved by discussion or by a third adjudicator. Single-reviewer screening is faster and demonstrably misses eligible studies.

Dual review pays off in the disagreements. Two reviewers who agree on everything are usually applying a criterion too loose to discriminate. A pilot round on fifty records, with the disagreement rate examined before full screening starts, is the cheapest quality control available on a review. Ambiguity in the criteria surfaces while there are still only fifty decisions to revisit, instead of three thousand.

Systematic review vs literature review

The comparison people search for most is the one most easily settled. A narrative literature review summarises what an author has read. For a newcomer to a field, a well-written narrative review by someone who knows that field often beats any systematic review; it makes no claim to completeness, and a reader has no way to check what was left out.

A systematic review makes exactly that claim and exposes the work that supports it. The trade is transparency for effort: months of process, in return for a result that survives scrutiny.

A narrative review vs systematic review question is usually a question about purpose. To know what a field thinks, read a narrative review. To know what the evidence supports, you need the systematic version. That is also what defends the answer to an editor, a guideline panel or a regulator.

When a systematic review is the right method

The method is expensive. It is worth it when three conditions hold together.

The question is specific enough to answer. Systematic reviews answer bounded questions: does this intervention affect this outcome in this population? Framed with PICO (population, intervention, comparator, outcome), a good question tells you immediately whether a study belongs. “What do we know about long COVID” is not a systematic review question. “Do graded exercise programmes improve fatigue scores in adults with post-viral fatigue” is.

The answer will be acted on. Systematic reviews underpin clinical guidelines, health technology assessments, reimbursement dossiers and policy decisions. Where nothing turns on the answer, the rigour is hard to justify.

Enough primary research exists. If three studies address your question, a systematic review will find three studies and say so. That is a legitimate and publishable finding. Consider whether it is the finding you need.

When a scoping or rapid review fits better

If the question is broad and you genuinely don’t know what has been published, you want a scoping review: it maps the literature where a systematic review estimates an effect, and its criteria may be refined as understanding develops. If you need an answer in weeks and a full review would take months, a rapid review applies systematic methods with documented shortcuts. Both are legitimate methods with standards of their own. Reporting one as the other is a problem.

Figure 1. Scoping review or systematic review: two questions that settle it.

Figure 2. Review types ordered by how much of the method is fixed before the first search runs.

The stages of a systematic literature review

Eight stages, from question to publication.

1 · Question and protocol.

Formulate the question, usually in PICO. Write the protocol: this includes the eligibility criteria, search strategy, screening process, extraction fields, appraisal tool, and analysis plan. This stage takes longer than people expect and saves more time than any other.

2 · Registration.

Register the protocol before screening begins. Registration timestamps your intentions, so a later deviation stays visible and can be explained. Journals increasingly ask for a registration number at submission.

3 · Search strategy.

Translate the question into search strings for each database: typically MEDLINE, Embase, CENTRAL, and subject-specific sources. Strategies are database-specific: a MEDLINE string does not run correctly on Embase. This is specialist work, and reviews with librarian involvement retrieve more relevant records.

4 · Deduplication.

Database exports overlap heavily. Deduplication typically removes twenty to forty per cent of a raw import. Doing it badly costs you twice: duplicates screened separately can receive contradictory decisions, and a record wrongly deleted as a duplicate is silently lost.

5 · Title and abstract screening.

Two reviewers, screening independently, against the protocol criteria. Volume is the problem. A broad clinical question routinely returns several thousand records after deduplication, most of them obviously irrelevant, which is what makes the work both necessary and tedious.

6 · Full-text review.

Retrieve full texts for records that pass screening and assess them against the same criteria. Every exclusion at this stage needs a recorded reason. PRISMA requires the breakdown, and reviewers check it. “Excluded: 143” without reasons is one of the most common causes of a manuscript being returned.

7 · Extraction and risk of bias.

Pull the study characteristics, methods, and results into a consistent structure, and assess each study for risk of bias with a tool matched to its design: RoB 2 for randomised trials, ROBINS-I for non-randomised interventions, QUADAS-2 for diagnostic accuracy, and so on. Extraction is normally done in duplicate as well.

8 · Synthesis and reporting.

Combine the findings, statistically where the studies allow it, and narratively where they don’t. Report against PRISMA 2020, including the flow diagram that accounts for every record from import to inclusion.

Figure 3. Eight stages of a systematic review, with a typical duration and an owning role for each.

How long does a systematic review take?

The literature puts a full systematic review at twelve to twenty-four months. Teams are often surprised by the distribution. Search and screening feel like the bulk of the work, and full-text retrieval, extraction, and appraisal reliably take longer.

Some stages compress, while others hold their length, no matter the time you spend on them.

Stages that respond to tooling. Deduplication is a solved problem and should be automatic. Screening can be reordered so that likely-relevant records surface first, which changes when you find the eligible studies but not how many records you assess. Retrieval of open-access full texts can be automated. Extraction can be templated, which is what keeps the data out of twelve differently formatted spreadsheets.

Stages that don’t. Formulating a good question is thinking. Judging whether a borderline study meets your criteria is judgement. Assessing risk of bias means reading the methods section and forming a view about it. Synthesis is interpretation.

Any claim that a systematic review can be produced in a week is describing something else. What the tooling actually changes is the ratio: less time on the mechanical parts, more available for the parts that need expertise. That is a smaller promise than the one usually made.

Who does what on a review team

Reviews are done by teams, and the roles are more distinct than they look from outside.

Information specialist or research librarian. Builds and documents the search strategy, advises on database selection, runs the searches, and handles the export and deduplication. In most institutions this person also brings review methodology into the team, which makes them the most consequential collaborator on the project and the one most often left out of planning.

Reviewers. Screen, retrieve, assess eligibility, and extract. Usually two, sometimes more on large reviews. They need to be calibrated against each other before volume screening begins. A pilot on fifty records surfaces disagreements about criteria while they are still cheap to fix.

Methodologist. Owns the protocol, the appraisal approach and the synthesis plan. On experienced teams this is one of the reviewers; on others it is a specialist brought in. Reviews without methodological input are where avoidable problems concentrate.

Statistician. Necessary when a meta-analysis is planned. Involve statisticians before extraction. The fields you need for pooling determine what the extraction form must capture, and retrofitting that is painful.

Project lead. Keeps the schedule, resolves adjudications, and handles the submission. On a two-person review this is typically one of the reviewers wearing a second hat.

Figure 4.Which role leads, contributes or stays out at each of the eight stages.

PRISMA guidelines for systematic reviews

PRISMA 2020 is the reporting standard for systematic reviews. It specifies a checklist of items your manuscript must address and a flow diagram accounting for every record from identification through to inclusion. Most journals in health research require it, and the PRISMA guidelines for systematic review reporting are now effectively a condition of submission.

The flow diagram is where reporting problems surface first, because the numbers have to reconcile. Records identified, minus duplicates removed, equals records screened. Records screened, minus records excluded, equals reports sought. When those chains break, a reviewer sees it immediately.

Related standards cover adjacent work: PRISMA-ScR for scoping reviews, PRISMA-P for protocols, PRISMA-S for search reporting.

Registration happens alongside all of these parts. PROSPERO covers systematic reviews with a health-related outcome. Scoping reviews go to the Open Science Framework or JBI. Register before any work begins. A registration filed after the fact records the outcome, and the point of registering was to record the intention.

Both Cochrane and NICE publish methods guidance that goes well beyond the reporting minimum, and teams working toward guideline or HTA submission should read the relevant handbook alongside the checklist.

Where reviews go wrong

Certain failures recur often enough to be worth naming. All of them are ordinary, and all of them are cheaper to prevent than to fix at peer review.

1 · Criteria that drift.

The protocol says one thing, the screening does another, and nobody writes down when the change happened. This usually starts innocently: an obviously relevant study fails a criterion, so the criterion widens. The fix is recording. Deviations from protocol are normal and acceptable once they are documented and justified. The trouble comes when a review’s final criteria trace back to nothing.

2 · Screening decisions without reasons.

At title and abstract stage a bare include/exclude is defensible, because the volume makes anything else impractical. At full-text stage the bar rises: PRISMA requires exclusion reasons broken down by category, and a reviewer who sees a single aggregate number will ask for the breakdown. Capturing the reason at the moment of the decision costs seconds. Reconstructing it three months later from a spreadsheet of PDF filenames costs days.

3 · Extraction into inconsistent structures.

Two reviewers extracting into their own spreadsheets will produce two schemas, and reconciling them is a substantial piece of work that nobody planned for. Worse, a statistician arriving after extraction routinely finds that the fields needed for pooling were never captured: outcome definitions recorded as prose where numbers were needed, standard deviations omitted because at the time nobody knew a meta-analysis was coming.

4 · Searches that cannot be re-run.

A search documented as “MEDLINE and Embase, 2015–2024” is not documented. Without the full strings, the field tags, and the date the search was executed, the review cannot be updated and its completeness cannot be checked. This matters more as living reviews become common: an undocumented search leaves a review that can only ever be redone from scratch.

5 · Numbers that don’t reconcile.

The PRISMA flow diagram is arithmetic, and the arithmetic is checkable. Records identified minus duplicates removed must equal records screened. When it fails, usually because deduplication was run twice or because records removed for other reasons were never counted, the discrepancy is visible to anyone who adds up the boxes.

Each of these is a failure of record-keeping, not of method. The reviewers knew what they were doing at the time. The trail of what they did was never captured in a form that survived the project.

Advantages and disadvantages of a systematic review

The advantages of a systematic review are real and worth stating plainly. The method reduces the influence of any single author’s reading. It makes the basis of a conclusion visible and checkable. It produces something that can be updated rather than redone, provided the search is documented well enough to re-run. And where enough comparable studies exist, it supports quantitative synthesis, which no narrative approach can offer.

The disadvantages are equally real. Systematic reviews are slow and expensive, and by the time one is published, the evidence base has usually moved. They can only summarise what has been published, so publication bias in that literature carries straight through into the result. Rigid criteria can exclude studies a specialist would consider informative. And the apparatus of the method (the flow diagram, the appraisal tables, the pooled estimate) carries an authority that the underlying evidence may not deserve. A well-conducted review of six poor trials is a well-conducted review of poor evidence, and readers routinely miss that distinction.

Living reviews are one response to obsolescence. The review is maintained, searches re-run on a schedule, new studies folded in as they appear. It changes the workflow from a project into a process, and it works only where the original search and screening decisions were documented well enough to extend.

Figure 5. Six things worth having in place before the first record is screened.

 


PICO Portal supports systematic review teams through screening, full-text review, extraction, risk-of-bias assessment and PRISMA-ready reporting, with reviewer decisions and their reasoning recorded throughout.

See how the workflow fits together.