AI for Researchers
Citation-grounded tools, deep-research agents, and how to check either one
{{Presenter Name}}
Landscape as of August 2026
2-Minute Recap
Today (Session 2): we live inside Tier 2 and Tier 3, and learn to verify what each one hands us.
Learning Objectives
Objectives 4–6 continue on the next slide.
Learning Objectives
Section 01
Two different matching methods, two different blind spots. Knowing which one you are using tells you what you have not seen.
Literacy Foundation
Literacy Foundation
Literacy Foundation
The Evidence
The Evidence
Section 02
Six citation-grounded tools, described from their own documentation — what each indexes, how each grounds an answer, and what it costs.
Tier 2 — Research-Specific Tools
Series Artefact
| Tool | Corpus (vendor-stated) | How it grounds an answer | Distinctive strength |
|---|---|---|---|
| Elicit | 125M–138M papers [6] [7] | Retrieval, then reports "inspired by systematic reviews" with sentence-level citations [6] | Structured extraction into tables across up to 1,000 papers [6]; full Research Agent since Aug 2026 [33] |
| Consensus | 200M+ documents [1] | Hybrid BM25 + embedding search, quality re-rank, then synthesis [1] | Fast question-shaped synthesis over peer-reviewed work [1] |
| Semantic Scholar | 214M papers, 2.49B citations [12] | Scholarly index, not a synthesiser; open API and SPECTER2 embeddings [12] | Free non-profit infrastructure many other tools are built on [11] |
| Scite | 280M+ full-text articles [13] | Smart Citations classify later work as supporting, contrasting or mentioning [13] | Licensed full text behind paywalls; 30+ publisher agreements [13] |
| ResearchRabbit | 310M+ papers [15] | Citation-graph expansion from seed papers; no synthesis layer [15] | Visual maps of how a literature connects and evolves [15] |
| SCiNiTO | 500M+ works, on OpenAlex [18] | Search with Boolean + filters, then AI chat that cites its sources [18] [19] | Open catalogue base plus journal recommender and PDF analysis [18] |
Series Artefact
| Tool | Citation export | Access / cost | Best fit — and its limit |
|---|---|---|---|
| Elicit | .bib and .ris; library itself not exportable [8] | Free Basic; Pro $49/user/mo (billed $588/yr) [7] | Extraction tables — but measured sensitivity was 39.5% (Feb–Mar 2025) [3] |
| Consensus | CSV and RIS into EndNote / Zotero / Mendeley [10] | Free tier; Premium $11.99/mo; 40% student discount [9] | Scoping a question fast — not an exhaustive search [1] |
| Semantic Scholar | Open API; most endpoints need no key [12] | Free [11] | Reliable metadata and citation links — no synthesis [11] |
| Scite | Zotero plugin, API, MCP into chat tools [13] | No free tier; Basic $20/mo, Pro $50/mo [14] | "Has this been contradicted?" — classification is automated [13] |
| ResearchRabbit | BibTeX out; Zotero import is one-way so far [17] | Free forever tier; RR+ from $10/mo [16] | Seed-paper expansion — needs a good seed [15] |
| SCiNiTO | Bookmarks and Spaces; export formats not documented [18] | Free individual tier; institutional unlocks full text [18] | Open-catalogue breadth — inherits OpenAlex's gaps [29] |
Critical Literacy
All corpus figures on this slide and on slides 13–14 are vendor self-reported, taken from each tool's own current pages; none is independently audited.
Section 03
They plan, browse, read and return a long report with citations. Three audits tell us exactly how far to trust that report.
Tier 3 — Deep-Research Agents
Tier 3 — Deep-Research Agents
The Evidence
Critical Literacy
Every performance figure in this deck comes from independent evaluation [3] [24] [25] [31] — no vendor self-reported performance claim appears on any slide.
Section 04
Systematic reviews, coverage bias, and paywalled literature — where these tools need more scrutiny than the interface suggests.
Caution Zone 1
Caution Zone 2
Caution Zone 3
Caution Zone 3
Series Artefact
Steps 1–2 catch fabrication; steps 3–5 catch misreading — now the dominant failure: with links valid, factual support runs 39–77% [31], and grounded tools do not remove it [1] [26].
Systematic Prompting
One answerable question, with the population, comparison and outcome named. Then constrain with the interface, not the prose.
Your field, what counts as evidence, a source per claim, and an explicit list of what it could not reach.
Live Demo & Takeaways
References
References