← series index

# Session 2: Literature Discovery & Synthesis

*AI for Researchers — attendee handout · Landscape as of August 2026*

---

## At a glance

| | |
|---|---|
| **Session** | 2 of 7 — Literature Discovery & Synthesis |
| **You'll learn** | Where semantic search, keyword search and deep-research agents each succeed and each fail, and how to verify what any of them hands you. |
| **You'll practise** | Running one real question through a citation-grounded tool and a deep-research agent, then checking three of the resulting citations. |
| **Prerequisite** | None. Session 1's tool taxonomy (Tier 1 chatbots / Tier 2 research tools / Tier 3 agents) is recapped in two minutes. |
| **Tools referenced** | Elicit, Consensus, Semantic Scholar, Scite, ResearchRabbit, SCiNiTO, and the deep-research agents — ChatGPT deep research, Gemini Deep Research / Deep Research Max, Claude Research — surveyed, not endorsed. |

---

## The one-paragraph version

Keyword search matches your words; semantic search matches your meaning. Each finds literature the other misses, and the best tools now run both and re-rank the result [1]. The catch is that ranking is not enumerating: in the sharpest independent peer-reviewed head-to-head — a **Feb–Mar 2025 measurement** of Elicit Pro's Review mode, a product Elicit has since replaced [3][33] — an AI search tool reached **39.5% sensitivity** against **94.5%** for the reviews' original Boolean searches, while reaching **41.8% precision** against **7.55%**, and turning up included studies the original searches had missed [3]. Deep-research agents go further, planning and browsing over several minutes to return a cited report [21][22][23]. A 2025 audit found citation accuracy of **40–80%** and persistent one-sidedness on contested questions [24]; a May-2026 preprint audit of **14 current-generation models** found the failure has moved, not gone: link validity now stays **above 94%**, yet factual support for the cited claims runs only **39–77%**, and fact-check accuracy dropped **~42%** as tool calls scaled from 2 to 150 [31]. Deep-research agents also hallucinate URLs at **10.7%** pooled versus **4.8%** for plain search-augmented models [25]. None of this makes the tools useless. It makes verification non-optional — and it moves the load-bearing check from "does the link resolve" to "does the source say that".

---

## The Literature-Discovery Tool Matrix

Every cell comes from the tool's own current documentation. Corpus figures are **vendor-stated and not independently audited**.

| Tool | Corpus (vendor-stated) | How it grounds an answer | Citation export | Access / cost | Best fit — and its limit |
|---|---|---|---|---|---|
| **Elicit** | 125M–138M papers [6][7] | Retrieval, then reports "inspired by systematic reviews", with sentence-level citations [6]; a full Research Agent (reports, data analysis, slides) since 2026-08-04 [33] | `.bib` and `.ris`; the library itself "cannot be exported at this time" [8] | Free Basic tier; Pro $49/user/mo (billed $588/yr); Scale $169/user/mo [7] | Structured extraction into comparison tables across up to 1,000 papers [6] — but measured sensitivity was 39.5% (Feb–Mar 2025 product; the Research Agent is unevaluated independently) [3][33] |
| **Consensus** | 200M+ scientific documents, from Semantic Scholar + OpenAlex + own crawl [1] | Hybrid BM25 + embedding search, quality re-rank of top 1,500, synthesis over top 20 (2025 description [1]); rebuilt in 2026 around a GPT-5-based multi-agent "Scholar Agent" [35] | CSV and RIS into EndNote / Zotero / Mendeley [10] | Free tier (unlimited search, 10 Pro Analyses/mo); Premium $11.99/mo; 40% student discount [9] | Scoping a question fast — not an exhaustive search [1] |
| **Semantic Scholar** | 214M papers, 2.49B citations, 79M authors [12] | A scholarly index rather than a synthesiser; open API with SPECTER2 embeddings [12] | Open API; most endpoints need no key [12] | Free; run by the non-profit Ai2 [11] | Reliable metadata and citation links, and the infrastructure other tools are built on [11] — no synthesis layer |
| **Scite** | 280M+ full-text articles, preprints, books, patents, datasets [13] | Smart Citations classify later work as supporting, contrasting or mentioning; answers link to the specific sentence [13] | Zotero plugin, API, MCP into ChatGPT/Claude [13] | No free tier; Basic $20/mo, Pro $50/mo, Team $50/seat/mo; 7-day trial [14] | Answering "has this finding been contradicted?" — via licensed full text from 30+ publishers [13]; classification is automated, so spot-check it |
| **ResearchRabbit** | 310M+ academic papers [15] | Citation-graph expansion from seed papers; no synthesis layer [15] | BibTeX out; the Zotero integration is import-only so far, with two-way sync announced [17] | Free Forever tier (up to 50 seed articles); RR+ from $10/mo; institutional tier with LibKey [16] | Seeing how a literature connects and evolves [15] — output quality depends entirely on your seed papers |
| **SCiNiTO** | 500M+ works, built on OpenAlex [18] | Search with Boolean AND/OR/NOT plus filters, then AI chat that cites real sources [18][19] | Bookmarks and shared Research Spaces; export formats not documented in the public help centre [18] | Guest (5 queries); Individual free tier with unlimited smart search; Institutional unlocks full text [18] | Open-catalogue breadth plus journal recommender and PDF analysis [18] — inherits OpenAlex's coverage shape [29] |

*Landscape as of August 2026; tool capabilities and pricing change quickly — verify before relying on this table. Row order is not a ranking.*

**Read the corpus column critically.** Elicit's home page says 125M papers while its own pricing page says 138M [6][7]. Scite's home page says 280M+ while its pricing page says 300M+, and the same home page describes its publisher agreements as both "30+" and "40+" — this table uses the conservative figure in each case [13][14]. SCiNiTO states 500M+ works on OpenAlex, while OpenAlex's own catalogue reports 243M works [18][20]. Nobody is lying — "works", "papers", "documents" and "sources" count different things. But a bigger index is not a better search.

---

## Prompt templates

Both templates are the Session 1 **CRIT** skeleton (Context / Role and register / Instructions and constraints / Task) applied to discovery.

**Template 1 — the discovery-tool query**

```
[One answerable question, naming the population, the comparison and the
outcome. One claim at a time.]

Example:
Can AI-based literature search tools achieve the same recall as traditional
Boolean database searching in systematic reviews?
```

Then constrain with the **interface**, not the prose: study type, year range, open access, journal filters.

*When to use:* whenever you want a checkable shortlist rather than an argument.
*Watch out for:* the tool re-ranks before you see anything — Consensus's 2025 pipeline filtered 200M documents down to 1,500, then to 20, applying recency, citation count and journal impact as quality signals [1]; its 2026 Scholar Agent rebuild changes the architecture, not the fact that something ranks before you see [35]. A correct, important, low-cited 2016 paper can be ranked out of your view.

**Template 2 — the deep-research agent brief**

```
Context: I am a researcher in [field] working on [specific question].
[Any background the agent could not guess.]

Role and register: write for [audience]. Neutral academic prose.

Instructions and constraints:
- Prioritise peer-reviewed work and official guidance over vendor material.
- Every quantitative claim must be followed by a source with a working link.
- Present the strongest case FOR and the strongest case AGAINST, and say
  which is better evidenced.
- List separately at the end: (a) sources you cited but could not read in
  full, and (b) anything you looked for and could not access.
- Do not use vendor blog posts as evidence of performance. If you cite one,
  label it a vendor claim.
- [Length limit], excluding the source list.

Task: [one specific request].
```

*When to use:* orientation in an unfamiliar debate, or scoping before a real search.
*Watch out for:* the "strongest case FOR / AGAINST" clause is doing heavy lifting. In the 2025 audit, deep-research configurations remained "highly one-sided on debate queries" [24] — no current-generation re-measurement of balance exists, so keep demanding the other side. And **edit the proposed research plan before it runs**: all three major agents let you [21][22], and almost nobody does.

Three things worth knowing about the agents themselves (August 2026): ChatGPT's deep research runs a GPT-5.2-based model — not the GPT-5.6 chat flagship — and its published per-tier quotas date to mid-2025, so check your own account [21][34]. Google's Deep Research and Deep Research Max both run on Gemini 3.1 Pro [32]. And the Tier 2 / Tier 3 boundary is blurring: Scite, Elicit and Consensus all expose their corpora to chatbots via MCP servers [13][33][36], so "which tier am I in?" is becoming "which corpus is this answer actually grounded in?"

---

## Checklist: verify an AI literature review

Run this on every citation you intend to keep. Stop at the first failure — you have already learned what you needed.

- [ ] **1. Resolve it.** Does the DOI or URL open a real record? *(In 2025-era audits 5–18% of citation URLs did not resolve and 3–13% were never real [25]. Current frontier systems keep link validity above 94% [31] — so a resolving link is now **necessary, not sufficient**. Do not stop here.)*
- [ ] **2. Match all four.** Author, year, journal and title against the record — not three of the four.
- [ ] **3. Open the full text.** Not the abstract. Roughly four in five catalogued works are not open access [20], so route it through your institutional access.
- [ ] **4. Find the sentence.** Locate the specific passage the claim rests on. If you cannot find it, the claim fails.
- [ ] **5. Compare the strength.** Does the source say *that*, or something weaker, narrower, or opposite? *(This is the failure citation-grounded retrieval does **not** remove [1] — and on current systems it is the dominant one: across 14 current-generation models with links valid, factual support ran only 39–77% [31]. The 2023 baseline pointed the same way: even real GPT-3.5/GPT-4 citations carried substantive errors in 24–43% of cases [26].)*
- [ ] **6. Check the direction of travel.** Has later work supported or contradicted it? A citation-context tool answers this in one click [13].
- [ ] **7. Check what is missing.** Which languages, regions and paywalled work could this tool not see? *(Of 62,701 active open-access journals, WoS indexes 6,157, Scopus 7,351, OpenAlex 34,217 — and only 4,094 appear in all three [29].)*
- [ ] **8. Check the benchmark before believing a performance claim.** Ask what counted as a hit and how many topics the number was averaged over. The same tool, in the same mode, evaluated by the same team over about two weeks, scored anywhere from 25.5% to 69.2% sensitivity depending on the review it was tested against — the widely quoted 39.5% is the mean of four case studies and describes none of them [3].
- [ ] **9. If this is an evidence synthesis:** can you demonstrate the AI use "will not compromise the methodological rigour or integrity" of the review, and have you transparently reported any AI use that makes or suggests a judgement? [27]

---

## The three caution zones

**1. Systematic reviews.** Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence issued a joint position statement in 2025: authors remain "ultimately responsible", AI "should be used with human oversight", and any AI use that "makes or suggests judgements should be fully and transparently reported" [27]. The peer-reviewed scoping evidence agrees that LLMs "should only be used with caution and under human supervision" [28].

**2. Coverage bias.** Every tool inherits its index's shape. Africa is underrepresented in WoS, Scopus and OpenAlex alike, and high-income economies are consistently over-represented in Scopus and WoS [29]. The open catalogue is not simply worse — OpenAlex's reference coverage is comparable to WoS and Scopus on shared publications [30] — but it is differently shaped. Ask which index, then ask who it leaves out.

**3. Paywalled literature.** OpenAlex catalogues 243M works, of which 48M are open access [20]. For the rest, most tools are reading metadata and an abstract — not the methods, the limitations or the sample description. Scite's answer is licensing agreements with 30+ publishers [13]; SCiNiTO's institutional tier is the same bargain [18]. Yours is your library.

**And a fourth, for the record:** Google Scholar is not a safe harbour. Across 120 systematic reviews its *coverage* was 97.2%, but its *recall* was 72.8% against 81.6% for Embase and MEDLINE combined, and only 46.4% of included studies sat inside the downloadable first 1,000 results [2].

---

## Further reading

- Lau, O. & Golder, S. (2025). *Comparison of Elicit AI and Traditional Literature Searching in Evidence Syntheses Using Four Case Studies*. Cochrane Evidence Synthesis and Methods — https://doi.org/10.1002/cesm.70050 — the sharpest independent, peer-reviewed head-to-head; a Feb–Mar 2025 measurement of a product Elicit has since replaced, which is exactly why its method matters more than its numbers. Source of the 39.5%/94.5% figures. Start here.
- Flemyng, E. et al. (2025). *Position statement on AI use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and CEE 2025* — https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.ED000178/full — two pages, six key messages, and the rules you will actually be held to.
- Venkit, P. N. et al. (2025). *DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence*. arXiv:2509.04499 — https://arxiv.org/abs/2509.04499 — the 2025 baseline audit behind the 40–80% citation-accuracy figure and the one-sidedness finding.
- Onweller, H. et al. (2026). *Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents*. Preprint, arXiv:2605.06635 — https://arxiv.org/abs/2605.06635 — the current-generation companion: 14 models, link validity above 94%, factual accuracy 39–77%, and fact-check accuracy dropping ~42% as tool calls scale. The paper behind "resolving is necessary, not sufficient".
- Bramer, W. M., Giustini, D. & Kramer, B. M. R. (2016). *Comparing the coverage, recall, and precision of searches for 120 systematic reviews in Embase, MEDLINE, and Google Scholar*. Systematic Reviews — https://pmc.ncbi.nlm.nih.gov/articles/PMC4772334/ — pre-dates all of this, and is still the clearest demonstration that coverage and recall are different things.
- Sahu, G., Charlin, L. & Pal, C. (2026). *Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth*. arXiv:2605.29234 — https://arxiv.org/abs/2605.29234 — read it for the uncomfortable half: humans were 2.5× more likely than the best AI re-ranker to cite a direct collaborator.
- Maddi, A., Maisonobe, M. & Boukacem-Zeghmouri, C. (2025). *Geographical and disciplinary coverage of open access journals: OpenAlex, Scopus, and WoS*. PLOS ONE — https://doi.org/10.1371/journal.pone.0320347 — the numbers behind "your index has a shape".

*Full annotated source list for this session: see `sources.md`.*

---

*AI for Researchers · Session 2: Literature Discovery & Synthesis · Landscape as of August 2026*