From random searching to a structured, transparent, reproducible evidence search — for a thesis, protocol, review, or manuscript.
Written & reviewed by the Bayyinah team · Last reviewed 2026-07-18 · Read it, watch it, or listen — then take the tools with you.
Most beginners think a literature search means writing a full sentence into PubMed, Google Scholar, or an AI tool and collecting the first few papers. That is not searching — that is guessing. A real literature search is a structured scientific process, and it is the foundation the whole project stands on.
Expanded, the same journey is: Define → Break → Expand → Build → Search → Refine → Screen → Extract → Map → Document. This lesson walks each step, using one running clinical question the whole way through.
A weak search doesn't just miss a few papers — it quietly damages everything downstream. You may claim a gap that isn't real, repeat a study already done, pick the wrong design or a weak outcome, miss standard definitions or a key systematic review, build a protocol on incomplete evidence, and collect months of supervisor and reviewer corrections.
A good search doesn't only find papers — it builds the logic of your project
What is known · what is uncertain · what is missing · which populations, outcomes and designs were used · what worked · what limitations remain · what gap is defensible · what the next study should do.
Guessing at novelty · duplicating work · choosing the wrong design · using an unvalidated outcome · missing key reviews · defending a gap that collapses under review.
We'll build one search from start to finish around a single question:
It's an observational, association question — so the natural frame is PECO:
PECO — the running question, deconstructed
| Element | Meaning | In our question |
|---|---|---|
| P | Population | Hospitalized adults with type 2 diabetes |
| E | Exposure | Documented hypoglycemia during admission |
| C | Comparator | No documented hypoglycemia |
| O | Outcome | Inpatient falls |
Before opening any database, know the whole road. Don't search the full question as a sentence — the question is for thinking; the strategy is for databases.
It defines the concepts. It is not itself the search string.
Here: diabetes · hypoglycemia · falls · hospital setting.
Every concept has many words — synonyms, abbreviations, spelling variants.
MeSH (PubMed) and Emtree (Embase) find papers even when authors use different words.
OR inside a concept, AND between concepts.
Usually PubMed first, then Embase, then Cochrane — plus others by topic.
Too many results → narrow. Too few → broaden. Never perfect on the first try.
Sort into reviews, original studies, methods, guidelines, outcome-definition papers.
Turn scattered papers into structured, comparable knowledge.
A search you cannot reproduce is weak science.
The single idea that separates a beginner search from a researcher's search: databases store two languages. Search both.
Words in titles, abstracts and keywords. Catch new articles not yet indexed, author-specific wording, abbreviations and recent terminology. Miss papers that use different words.
The database's official index — MeSH in PubMed, Emtree in Embase. Groups "low blood sugar" and "hypoglycaemic event" together. Misses very new, not-yet-indexed articles.
Two more precision tools: phrase searching with quotes ("type 2 diabetes") and truncation with an asterisk (fall* → fall, falls, falling, fall-related). Both sharpen precision but can cost sensitivity — a search for only "inpatient falls" misses "falls among hospitalized patients." Field tags aim the search: in PubMed [tiab] = title/abstract, [MeSH] = subject heading.
OR inside each concept; AND between them. This is the logic — you translate it into each database next.
Concept blocks for the running question
"type 2 diabetes" OR T2DM OR "diabetes mellitus"hypoglycemia OR "low blood glucose" OR "low blood sugar"falls OR "accidental falls" OR "fall risk" OR "inpatient falls"Do not copy-paste blindly between databases — translate the logic, not just the words. Here is the same search, three ways.
Cochrane searches are simpler and concept-focused: diabetes AND hypoglycemia AND falls, or for interventions, diabetes AND fall prevention. For a pure association question, PubMed and Embase carry more weight — but Cochrane still surfaces fall-prevention and hypoglycemia-prevention trials and any existing reviews.
Start here — free, biomedical
Broaden — drugs & devices
Trials & reviews
Add as needed
Thesis background → PubMed + key reviewsProtocol → PubMed + Embase + CochraneSystematic review → all + registries + grey litDrug safety → Embase-led
Add a concept or the setting block · use [tiab] fields · use phrases · use more specific outcome terms · add an age/human filter · limit by design only when justified.
Remove a concept or the setting · add synonyms · drop exact phrases & filters · broaden terms · check spelling & controlled vocabulary · try another database · chase citations.
Date, language, humans, age, publication type, study design and full-text filters can sharpen a search — or silently bury the paper you needed. Use them only when justified.
For a thesis, review or manuscript, a search that isn't documented cannot be judged or repeated. Keep a simple search log.
Minimum search log
| Database | Date | Search string | Filters | Results | Notes |
|---|---|---|---|---|---|
| PubMed | [date] | full string | humans, adults | 325 | MeSH + tiab |
| Embase | [date] | full string | none | 610 | incl. conference abstracts |
| Cochrane CENTRAL | [date] | simplified | trials | 42 | intervention evidence |
Also record: platform, controlled vocabulary used, export file, deduplication method & count, searcher name.
Screen in stages — title → abstract → full text → final inclusion — against clear inclusion/exclusion criteria. Tools like Rayyan, Covidence, Zotero or EndNote manage this; for a thesis, a tidy spreadsheet can be enough. Then convert papers into a matrix — the point where a search becomes understanding.
Literature matrix columns
AuthorYearCountryDesignPopulationSample sizeExposure/interventionComparatorOutcome + definitionMain result + effect sizeLimitationsRisk-of-bias concernsRelevance to my questionGap identified
For our example, capture the details that decide the answer: how hypoglycemia was defined (threshold, symptomatic vs biochemical), the fall definition, timing of the fall relative to hypoglycemia, and adjusted confounders — age, frailty, renal disease, insulin use, polypharmacy.
Databases find records; citation chasing finds the conversation around a key paper. Read the reference list (backward) and the papers that cited it (forward), using Google Scholar, Scopus, Web of Science, Semantic Scholar, Research Rabbit, Connected Papers or Litmaps.
AI can genuinely accelerate a literature search — as an assistant, organiser, synonym generator, extractor, and reviewer of your logic. It cannot be the source of truth.
The AI tool ladder — exploration first, evidence last
| Tool type | Good for | Watch out for |
|---|---|---|
| General LLMs (ChatGPT, Gemini, Claude) | Break question into concepts, synonyms, draft Boolean, explain terms, matrix templates | Hallucinated references, wrong MeSH/Emtree — mark vocabulary "to verify" |
| Academic search (Elicit, Consensus, Perplexity, AnswerThis) | Orientation, related papers, common outcomes, evidence direction | Incomplete coverage, ranking bias — triage only, then open papers |
| Citation maps (Research Rabbit, Connected Papers, Litmaps) | Forward/backward chasing, finding clusters & seminal papers | Not a full search; bias toward highly-cited work |
| Screening (Rayyan, Covidence, ASReview) | Deduplication, title/abstract screening, PRISMA flow | Criteria must be human-defined; check disagreements |
| Extraction (Elicit, NotebookLM, SciSpace) | Pull PICO/PECO, design, outcomes, limitations into tables | Verify every extraction against the PDF |
A safe AI workflow — in order
Starter prompt: "Act as a clinical research mentor and medical librarian. My question is: [insert]. Break it into searchable concepts (population, exposure, comparator, outcome, setting). Don't search yet."
Quick check
Which is the safest way to use AI in a literature search?
A structured scientific process — not typing a sentence into PubMed. You move from question, to concepts, to synonyms and controlled vocabulary, to the right databases, then refine, document, and organise the papers into a usable evidence matrix.
Both. Free text catches new and author-specific wording; controlled vocabulary (MeSH, Emtree) catches papers indexed under different words. Together they give the best coverage.
Start with PubMed/MEDLINE; add Embase for drugs, adverse events, devices, international and conference coverage; use Cochrane/CENTRAL for trials and reviews. Add CINAHL, PsycINFO, Scopus, Web of Science, registries and grey literature by topic.
Use it to break down the question, generate synonyms, draft Boolean logic and suggest vocabulary to verify. Never cite AI output or an unopened paper; verify every reference and term; keep a prompt log.
The databases and tools named here are real, independently published resources — verify each in the tool itself.
Grab the free research tools, or get expert eyes on your search strategy, gap, and evidence matrix.