Bayyinah Research Academy · Volume 1 · Lecture 3

Literature Search Techniques

From random searching to a structured, transparent, reproducible evidence search — for a thesis, protocol, review, or manuscript.

Written & reviewed by the Bayyinah team · Last reviewed 2026-07-18 · Read it, watch it, or listen — then take the tools with you.

Most beginners think a literature search means writing a full sentence into PubMed, Google Scholar, or an AI tool and collecting the first few papers. That is not searching — that is guessing. A real literature search is a structured scientific process, and it is the foundation the whole project stands on.

Do not search harder. Search smarter. The Bayyinah search pathway:
Question Concepts Terms Databases Papers Matrix Gap

Expanded, the same journey is: Define → Break → Expand → Build → Search → Refine → Screen → Extract → Map → Document. This lesson walks each step, using one running clinical question the whole way through.

1. Why a weak search is dangerous

A weak search doesn't just miss a few papers — it quietly damages everything downstream. You may claim a gap that isn't real, repeat a study already done, pick the wrong design or a weak outcome, miss standard definitions or a key systematic review, build a protocol on incomplete evidence, and collect months of supervisor and reviewer corrections.

A good search doesn't only find papers — it builds the logic of your project

A strong search tells you

What is known · what is uncertain · what is missing · which populations, outcomes and designs were used · what worked · what limitations remain · what gap is defensible · what the next study should do.

A weak search leaves you

Guessing at novelty · duplicating work · choosing the wrong design · using an unvalidated outcome · missing key reviews · defending a gap that collapses under review.

2. Our running example

We'll build one search from start to finish around a single question:

Among hospitalized adults with type 2 diabetes, is documented hypoglycemia during admission associated with inpatient falls?

It's an observational, association question — so the natural frame is PECO:

PECO — the running question, deconstructed

ElementMeaningIn our question
PPopulationHospitalized adults with type 2 diabetes
EExposureDocumented hypoglycemia during admission
CComparatorNo documented hypoglycemia
OOutcomeInpatient falls

3. The ten-step journey

Before opening any database, know the whole road. Don't search the full question as a sentence — the question is for thinking; the strategy is for databases.

Start with the question

It defines the concepts. It is not itself the search string.

Break it into concepts

Here: diabetes · hypoglycemia · falls · hospital setting.

Expand each concept into terms

Every concept has many words — synonyms, abbreviations, spelling variants.

Add controlled vocabulary

MeSH (PubMed) and Emtree (Embase) find papers even when authors use different words.

Build search blocks

OR inside a concept, AND between concepts.

Choose the right databases

Usually PubMed first, then Embase, then Cochrane — plus others by topic.

Search, test, and refine

Too many results → narrow. Too few → broaden. Never perfect on the first try.

Screen and organise

Sort into reviews, original studies, methods, guidelines, outcome-definition papers.

Build a literature matrix

Turn scattered papers into structured, comparable knowledge.

Document everything

A search you cannot reproduce is weak science.

4. Free text vs controlled vocabulary

The single idea that separates a beginner search from a researcher's search: databases store two languages. Search both.

Free-text terms

Words in titles, abstracts and keywords. Catch new articles not yet indexed, author-specific wording, abbreviations and recent terminology. Miss papers that use different words.

Controlled vocabulary

The database's official index — MeSH in PubMed, Emtree in Embase. Groups "low blood sugar" and "hypoglycaemic event" together. Misses very new, not-yet-indexed articles.

Free text catches author language. Controlled vocabulary catches database language. Strong searches use both.

5. Boolean logic — the grammar of searching

The three Boolean operators — how each shapes your search
AND both narrows — both must appear diabetes AND foot ulcer OR widens — either counts diabetic OR mellitus NOT × excludes — removes a set foot NOT mouth

Two more precision tools: phrase searching with quotes ("type 2 diabetes") and truncation with an asterisk (fall* → fall, falls, falling, fall-related). Both sharpen precision but can cost sensitivity — a search for only "inpatient falls" misses "falls among hospitalized patients." Field tags aim the search: in PubMed [tiab] = title/abstract, [MeSH] = subject heading.

6. Build the blocks

OR inside each concept; AND between them. This is the logic — you translate it into each database next.

Concept blocks for the running question

Diabetes"type 2 diabetes" OR T2DM OR "diabetes mellitus"
AND
Hypoglycemiahypoglycemia OR "low blood glucose" OR "low blood sugar"
AND
Fallsfalls OR "accidental falls" OR "fall risk" OR "inpatient falls"
Add the hospital-setting block only if results are too broad. Adding too many concepts too early hides relevant studies.

7. Translate for each database

Do not copy-paste blindly between databases — translate the logic, not just the words. Here is the same search, three ways.

PubMed / MEDLINE — the biomedical starting point

("Diabetes Mellitus, Type 2"[MeSH] OR "type 2 diabetes"[tiab] OR T2DM[tiab])
AND
("Hypoglycemia"[MeSH] OR hypoglycemia[tiab] OR "low blood glucose"[tiab])
AND
("Accidental Falls"[MeSH] OR falls[tiab] OR "inpatient falls"[tiab] OR "hospital falls"[tiab])

Embase — broader drugs, adverse events, devices, conferences

'diabetes mellitus type 2'/exp OR 'type 2 diabetes':ti,ab OR T2DM:ti,ab
AND
'hypoglycemia'/exp OR hypoglycemia:ti,ab OR 'low blood glucose':ti,ab
AND
'accidental fall'/exp OR falls:ti,ab OR 'inpatient falls':ti,ab
Embase is not PubMed with different colours. It has its own vocabulary (Emtree), strengths, and syntax. Verify Emtree terms in the database.

Cochrane / CENTRAL — trials & systematic reviews

Cochrane searches are simpler and concept-focused: diabetes AND hypoglycemia AND falls, or for interventions, diabetes AND fall prevention. For a pure association question, PubMed and Embase carry more weight — but Cochrane still surfaces fall-prevention and hypoglycemia-prevention trials and any existing reviews.

8. Which databases, by purpose

PubMed / MEDLINE

Start here — free, biomedical

  • Clinical medicine, epidemiology, public health
  • MeSH indexing, strong filters
  • Best first stop for most clinicians

Embase

Broaden — drugs & devices

  • Pharmacology, adverse events, pharmacovigilance
  • International journals + conference abstracts
  • Usually needs a subscription

Cochrane / CENTRAL

Trials & reviews

  • Randomised & quasi-randomised trials
  • Existing systematic reviews, certainty of evidence
  • Key for intervention questions

Also, by topic

Add as needed

  • CINAHL (nursing/allied), PsycINFO (mental health)
  • Scopus / Web of Science (citation tracking)
  • ClinicalTrials.gov / WHO ICTRP, grey literature

Thesis background → PubMed + key reviewsProtocol → PubMed + Embase + CochraneSystematic review → all + registries + grey litDrug safety → Embase-led

9. Too many or too few results

Too many → focus

Add a concept or the setting block · use [tiab] fields · use phrases · use more specific outcome terms · add an age/human filter · limit by design only when justified.

Too few → expand

Remove a concept or the setting · add synonyms · drop exact phrases & filters · broaden terms · check spelling & controlled vocabulary · try another database · chase citations.

Too many results means focus. Too few results means expand. Both are normal — the first search is never the final search.

10. Filters: useful, but dangerous

Date, language, humans, age, publication type, study design and full-text filters can sharpen a search — or silently bury the paper you needed. Use them only when justified.

Never use "free full text" as a quality filter. A paper is not better because it is free, nor worse because it is behind a paywall. Filters are tools, not shortcuts.

11. Document it — or it isn't reproducible

For a thesis, review or manuscript, a search that isn't documented cannot be judged or repeated. Keep a simple search log.

Minimum search log

DatabaseDateSearch stringFiltersResultsNotes
PubMed[date]full stringhumans, adults325MeSH + tiab
Embase[date]full stringnone610incl. conference abstracts
Cochrane CENTRAL[date]simplifiedtrials42intervention evidence

Also record: platform, controlled vocabulary used, export file, deduplication method & count, searcher name.

If you cannot show how you searched, your reader cannot judge what you may have missed.

12. Screen, then build a literature matrix

Screen in stages — title → abstract → full text → final inclusion — against clear inclusion/exclusion criteria. Tools like Rayyan, Covidence, Zotero or EndNote manage this; for a thesis, a tidy spreadsheet can be enough. Then convert papers into a matrix — the point where a search becomes understanding.

Literature matrix columns

AuthorYearCountryDesignPopulationSample sizeExposure/interventionComparatorOutcome + definitionMain result + effect sizeLimitationsRisk-of-bias concernsRelevance to my questionGap identified

For our example, capture the details that decide the answer: how hypoglycemia was defined (threshold, symptomatic vs biochemical), the fall definition, timing of the fall relative to hypoglycemia, and adjusted confounders — age, frailty, renal disease, insulin use, polypharmacy.

13. Citation chasing — find the conversation

Databases find records; citation chasing finds the conversation around a key paper. Read the reference list (backward) and the papers that cited it (forward), using Google Scholar, Scopus, Web of Science, Semantic Scholar, Research Rabbit, Connected Papers or Litmaps.

14. Using AI safely

AI can genuinely accelerate a literature search — as an assistant, organiser, synonym generator, extractor, and reviewer of your logic. It cannot be the source of truth.

The AI tool ladder — exploration first, evidence last

Tool typeGood forWatch out for
General LLMs (ChatGPT, Gemini, Claude)Break question into concepts, synonyms, draft Boolean, explain terms, matrix templatesHallucinated references, wrong MeSH/Emtree — mark vocabulary "to verify"
Academic search (Elicit, Consensus, Perplexity, AnswerThis)Orientation, related papers, common outcomes, evidence directionIncomplete coverage, ranking bias — triage only, then open papers
Citation maps (Research Rabbit, Connected Papers, Litmaps)Forward/backward chasing, finding clusters & seminal papersNot a full search; bias toward highly-cited work
Screening (Rayyan, Covidence, ASReview)Deduplication, title/abstract screening, PRISMA flowCriteria must be human-defined; check disagreements
Extraction (Elicit, NotebookLM, SciSpace)Pull PICO/PECO, design, outcomes, limitations into tablesVerify every extraction against the PDF

A safe AI workflow — in order

Break question Synonyms Vocab (verify) Boolean blocks Translate Test in real DBs Map + extract Verify papers

Starter prompt: "Act as a clinical research mentor and medical librarian. My question is: [insert]. Break it into searchable concepts (population, exposure, comparator, outcome, setting). Don't search yet."

AI safety rules: never cite an AI answer as evidence · never cite a paper you didn't open · verify title, authors, journal, year, DOI/PMID · verify every MeSH/Emtree term and every database's syntax · never upload identifiable patient data · keep a prompt log. AI is your co-pilot, not your principal investigator.

15. Quality checklist

  • Is the research question clear, and broken into searchable concepts?
  • Did I gather synonyms, and check MeSH (and Emtree, if using Embase)?
  • Did I combine OR and AND correctly, and search more than one database when needed?
  • Did I use filters carefully, and test that known key papers are retrieved?
  • Did I export, deduplicate, and record the date and full search string?
  • Did I build a matrix, verify AI-assisted outputs manually, and could someone else reproduce my search?
Search like a researcher, not like someone scrolling for papers. A structured search saves time later because it prevents weak decisions early.

Quick check

Which is the safest way to use AI in a literature search?

16. Frequently asked questions

What is a proper literature search?

A structured scientific process — not typing a sentence into PubMed. You move from question, to concepts, to synonyms and controlled vocabulary, to the right databases, then refine, document, and organise the papers into a usable evidence matrix.

Free text or controlled vocabulary — which should I use?

Both. Free text catches new and author-specific wording; controlled vocabulary (MeSH, Emtree) catches papers indexed under different words. Together they give the best coverage.

Which databases should I search?

Start with PubMed/MEDLINE; add Embase for drugs, adverse events, devices, international and conference coverage; use Cochrane/CENTRAL for trials and reviews. Add CINAHL, PsycINFO, Scopus, Web of Science, registries and grey literature by topic.

How do I use AI without damaging the science?

Use it to break down the question, generate synonyms, draft Boolean logic and suggest vocabulary to verify. Never cite AI output or an unopened paper; verify every reference and term; keep a prompt log.

Sources & further reading

The databases and tools named here are real, independently published resources — verify each in the tool itself.

Turn a strong search into a strong thesis or manuscript

Grab the free research tools, or get expert eyes on your search strategy, gap, and evidence matrix.