AI for Researchers

Session 2: Literature Discovery & Synthesis

Citation-grounded tools, deep-research agents, and how to check either one

{{Presenter Name}}

Landscape as of August 2026

2-Minute Recap

The Series So Far


Today (Session 2): we live inside Tier 2 and Tier 3, and learn to verify what each one hands us.

Learning Objectives

By the End of This Session You Will Be Able To…


  1. Explain the difference between semantic and keyword search and why each surfaces literature the other misses.
  2. Compare citation-grounded discovery tools (e.g., Elicit, Consensus, Semantic Scholar, SCiNiTO, Scite, ResearchRabbit) on coverage, grounding, and fit for a given task.
  3. Run a literature search on the same question using both a citation-grounded discovery tool and a deep-research agent.

Objectives 4–6 continue on the next slide.

Learning Objectives

…And You Will Also Be Able To


  1. Evaluate deep-research agent output for coverage bias, missing paywalled literature, and unverified claims.
  2. Apply a citation-verification workflow to confirm a claimed source actually supports the claim made from it.
  3. Identify caution zones — systematic reviews, coverage bias, paywalled literature — where AI literature tools need extra scrutiny.

Section 01

How Search Finds — and Misses — Papers

Two different matching methods, two different blind spots. Knowing which one you are using tells you what you have not seen.

Literacy Foundation

Two Ways a Search Engine Can Match


Keyword (lexical)

  • Matches the words you typed.
  • Classic method: BM25. [1]
  • Boolean AND / OR / NOT, field limits, exact phrases.
  • Deterministic: same query, same results, next year.

Semantic (embedding)

  • Matches the meaning of your question.
  • Query and papers become vectors; nearest ones win. [1]
  • Finds synonyms and paraphrases you never thought of.
  • Ranked, opaque, and it changes when the model changes.

Literacy Foundation

The Good Tools Run Both, Then Re-Rank


Literacy Foundation

What Each One Misses


Keyword misses…

  • Papers using a different vocabulary for your concept.
  • Adjacent fields that named the same idea differently.
  • Anything past the result cap — only 46.4% of included studies sat in Google Scholar's first 1,000 hits. [2]

Semantic misses…

  • Exhaustiveness — it ranks, it does not enumerate.
  • Precise Boolean control over population, design or outcome.
  • Reproducibility — the ranking is not stable across model updates.

The Evidence

Measured Head-to-Head: Sensitivity vs. Precision


39.5% vs 94.5% Sensitivity — Elicit Pro vs. the original searches (Feb–Mar 2025) [3]
41.8% vs 7.55% Precision — Elicit Pro vs. the original searches (Feb–Mar 2025) [3]

The Evidence

Before We Get Smug: Human Search Is Biased Too


Section 02

The Discovery Toolkit

Six citation-grounded tools, described from their own documentation — what each indexes, how each grounds an answer, and what it costs.

Tier 2 — Research-Specific Tools

What "Citation-Grounded" Actually Buys You


Series Artefact

The Literature-Discovery Tool Matrix (1 of 2): Corpus & Grounding


Tool Corpus (vendor-stated) How it grounds an answer Distinctive strength
Elicit 125M–138M papers [6] [7] Retrieval, then reports "inspired by systematic reviews" with sentence-level citations [6] Structured extraction into tables across up to 1,000 papers [6]; full Research Agent since Aug 2026 [33]
Consensus 200M+ documents [1] Hybrid BM25 + embedding search, quality re-rank, then synthesis [1] Fast question-shaped synthesis over peer-reviewed work [1]
Semantic Scholar 214M papers, 2.49B citations [12] Scholarly index, not a synthesiser; open API and SPECTER2 embeddings [12] Free non-profit infrastructure many other tools are built on [11]
Scite 280M+ full-text articles [13] Smart Citations classify later work as supporting, contrasting or mentioning [13] Licensed full text behind paywalls; 30+ publisher agreements [13]
ResearchRabbit 310M+ papers [15] Citation-graph expansion from seed papers; no synthesis layer [15] Visual maps of how a literature connects and evolves [15]
SCiNiTO 500M+ works, on OpenAlex [18] Search with Boolean + filters, then AI chat that cites its sources [18] [19] Open catalogue base plus journal recommender and PDF analysis [18]
Landscape as of August 2026. Every cell is taken from the tool's own current documentation; corpus figures are vendor-stated and not independently audited. Row order is not a ranking — no endorsement implied.

Series Artefact

The Literature-Discovery Tool Matrix (2 of 2): Export, Cost & Fit


Tool Citation export Access / cost Best fit — and its limit
Elicit .bib and .ris; library itself not exportable [8] Free Basic; Pro $49/user/mo (billed $588/yr) [7] Extraction tables — but measured sensitivity was 39.5% (Feb–Mar 2025) [3]
Consensus CSV and RIS into EndNote / Zotero / Mendeley [10] Free tier; Premium $11.99/mo; 40% student discount [9] Scoping a question fast — not an exhaustive search [1]
Semantic Scholar Open API; most endpoints need no key [12] Free [11] Reliable metadata and citation links — no synthesis [11]
Scite Zotero plugin, API, MCP into chat tools [13] No free tier; Basic $20/mo, Pro $50/mo [14] "Has this been contradicted?" — classification is automated [13]
ResearchRabbit BibTeX out; Zotero import is one-way so far [17] Free forever tier; RR+ from $10/mo [16] Seed-paper expansion — needs a good seed [15]
SCiNiTO Bookmarks and Spaces; export formats not documented [18] Free individual tier; institutional unlocks full text [18] Open-catalogue breadth — inherits OpenAlex's gaps [29]
Landscape as of August 2026. Pricing and features change monthly — re-check before relying on a cell. Where a vendor does not document a capability, this table says so rather than guessing.

Critical Literacy

Read the Corpus Number Critically


All corpus figures on this slide and on slides 13–14 are vendor self-reported, taken from each tool's own current pages; none is independently audited.

Section 03

Tier 3: Deep-Research Agents

They plan, browse, read and return a long report with citations. Three audits tell us exactly how far to trust that report.

Tier 3 — Deep-Research Agents

What the Agent Actually Does


Tier 3 — Deep-Research Agents

Where Agents Genuinely Win


The Evidence

Documented Failure Modes


Critical Literacy

The Same Tool Scored 25% and 69%


Every performance figure in this deck comes from independent evaluation [3] [24] [25] [31] — no vendor self-reported performance claim appears on any slide.

Section 04

Three Caution Zones

Systematic reviews, coverage bias, and paywalled literature — where these tools need more scrutiny than the interface suggests.

Caution Zone 1

Systematic Reviews Have Rules — and They Now Cover AI


Caution Zone 2

Coverage Bias: The Index Shapes the Answer


Caution Zone 3

Paywalls: Most Tools Are Reading Abstracts


Caution Zone 3

And No, Google Scholar Is Not the Safe Harbour


Series Artefact

The Citation-Verification Workflow


  1. Resolve it. Necessary, no longer sufficient: frontier systems now keep link validity above 94% [31] — in 2025 audits 5–18% still failed here. [25]
  2. Match it. Do authors, year, journal and title agree with the record — all four?
  3. Open it. Find the sentence in the source that carries the claim.
  4. Compare it. Does the source say that, or something weaker, narrower or opposite? [24] [31]
  5. Check the direction. Was the paper later contradicted? [13]

Steps 1–2 catch fabrication; steps 3–5 catch misreading — now the dominant failure: with links valid, factual support runs 39–77% [31], and grounded tools do not remove it [1] [26].

Systematic Prompting

CRIT, Applied to Discovery


A discovery tool

One answerable question, with the population, comparison and outcome named. Then constrain with the interface, not the prose.

  • C: the question, in one sentence.
  • I: year, design and open-access filters.
  • T: one claim at a time — then re-ask.

A deep-research agent

Your field, what counts as evidence, a source per claim, and an explicit list of what it could not reach.

  • C+R: disciplines, audience, register.
  • I: demand disconfirming evidence — audited systems did not volunteer it. [24]
  • T: edit the plan before it runs. [21]

Live Demo & Takeaways

What We Do Next — and What You Take Home


References

References (1–18)


  1. Consensus (2025). Welcome to Consensus. consensus.app/home/blog/welcome-to-consensus/ — accessed 2026-07-27
  2. Bramer, Giustini & Kramer (2016). Comparing the coverage, recall, and precision of searches for 120 systematic reviews… Systematic Reviews 5, 39. pmc.ncbi.nlm.nih.gov/articles/PMC4772334/ — accessed 2026-07-27
  3. Lau & Golder (2025). Comparison of Elicit AI and Traditional Literature Searching in Evidence Syntheses Using Four Case Studies. Cochrane Evidence Synthesis and Methods 3(5), e70050. doi.org/10.1002/cesm.70050 — accessed 2026-07-27
  4. Sahu, Charlin & Pal (2026). Rethinking Literature Search Evaluation… arXiv:2605.29234. arxiv.org/abs/2605.29234 — accessed 2026-07-27
  5. Pan et al. (2026). ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search… arXiv:2606.20235. arxiv.org/abs/2606.20235 — accessed 2026-07-27
  6. Elicit (2026). Elicit: AI for scientific research. elicit.com/ — accessed 2026-07-27
  7. Elicit (2026). Pricing | Elicit. elicit.com/pricing — accessed 2026-07-27
  8. Elicit (2026). Export your data from Elicit. support.elicit.com/en/articles/1153857 — accessed 2026-07-27
  9. Consensus (2026). Pricing — Consensus. consensus.app/home/pricing/ — accessed 2026-07-27
  10. Consensus (2026). How to Export Consensus Results to Reference Managers. help.consensus.app/en/articles/9922811 — accessed 2026-07-27
  11. Semantic Scholar / Ai2 (2026). About Semantic Scholar. semanticscholar.org/about — accessed 2026-07-27
  12. Semantic Scholar (2026). Semantic Scholar Academic Graph API — Overview. semanticscholar.org/product/api — accessed 2026-07-27
  13. Scite (2026). AI for Research | Scite. scite.ai/ — accessed 2026-07-27
  14. Scite (2026). Scite Pricing. scite.ai/pricing — accessed 2026-07-27
  15. ResearchRabbit (2026). ResearchRabbit. researchrabbit.ai/ — accessed 2026-07-27
  16. ResearchRabbit (2026). Pricing | ResearchRabbit. researchrabbit.ai/pricing — accessed 2026-07-27
  17. ResearchRabbit (2026). Using the Zotero Importer. learn.researchrabbit.ai/en/articles/12796541 — accessed 2026-07-27
  18. SCiNiTO (2026). What is SCiNiTO? scinito.ai/help/getting-started/what-is-scinito — accessed 2026-07-27

References

References (19–36)


  1. SCiNiTO (2026). Boolean operators (AND, OR, NOT). scinito.ai/help/smart-search/boolean-operators — accessed 2026-07-27
  2. OpenAlex (2026). About OpenAlex. openalex.org/about — accessed 2026-07-27
  3. OpenAI (2026). Deep research in ChatGPT. help.openai.com/en/articles/10500283 — accessed 2026-07-27
  4. Google (2026). Gemini Deep Research Agent — Gemini API. ai.google.dev/gemini-api/docs/interactions/deep-research — accessed 2026-07-27
  5. Anthropic (2026). Using Research on Claude. support.claude.com/en/articles/11088861 — accessed 2026-07-27
  6. Venkit et al. (2025). DeepTRACE: Auditing Deep Research AI Systems… arXiv:2509.04499. arxiv.org/abs/2509.04499 — accessed 2026-07-27
  7. Rao et al. (2026). Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents. arXiv:2604.03173. arxiv.org/abs/2604.03173 — accessed 2026-09-02
  8. Walters & Wilder (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13, 14045. nature.com/articles/s41598-023-41032-5 — accessed 2026-07-27
  9. Flemyng et al. (2025). Position statement on AI use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and CEE 2025. Cochrane Database Syst Rev, ED000178. cochranelibrary.com/cdsr/doi/10.1002/14651858.ED000178/full — accessed 2026-07-27
  10. Lieberum et al. (2025). Large language models for conducting systematic reviews: on the rise, but not yet ready for use. J Clin Epidemiol 181, 111746. jclinepi.com/article/S0895-4356(25)00079-4/fulltext — accessed 2026-07-27
  11. Maddi et al. (2025). Geographical and disciplinary coverage of open access journals: OpenAlex, Scopus, and WoS. PLOS ONE 20(3), e0320347. doi.org/10.1371/journal.pone.0320347 — accessed 2026-07-27
  12. Culbert et al. (2025). Reference coverage analysis of OpenAlex compared to Web of Science and Scopus. Scientometrics. doi.org/10.1007/s11192-025-05293-3 — accessed 2026-07-27
  13. Onweller, H., et al. (2026). Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents. Preprint, arXiv:2605.06635. arxiv.org/abs/2605.06635 — accessed 2026-09-02
  14. Google (2026). Next-generation Gemini Deep Research. blog.google/…/next-generation-gemini-deep-research/ — accessed 2026-09-02
  15. Elicit (2026). Elicit Blog — Research Agent and API/MCP release notes. elicit.com/blog — accessed 2026-09-02
  16. Wikipedia (2026). Deep research (secondary; cites OpenAI documentation). en.wikipedia.org/wiki/Deep_research — accessed 2026-09-02
  17. OpenAI (2026). Consensus: building Scholar Agent on GPT-5 (case study). openai.com/index/consensus/ — accessed 2026-09-02
  18. Consensus (2026). Consensus MCP Server (documentation). docs.consensus.app/docs/mcp — accessed 2026-09-02