← series index
# Sources — Session 2: Literature Discovery & Synthesis
*AI for Researchers · dated, annotated reference list · compiled 2026-07-27 · revised 2026-09-02 · Landscape as of August 2026*
Entry format (one line per source, numbering matches the `[n]` footnote markers used in `slides.html` and `handout.md`):
`- [n] Author/Org (Year). *Title*. URL — accessed YYYY-MM-DD. [type: peer-reviewed | policy | primary-doc | news] — one-line note on what claim(s) it grounds.`
Allowed `type` values:
- `peer-reviewed` — peer-reviewed paper or arXiv/preprint literature
- `policy` — official publisher/funder/institutional policy page
- `primary-doc` — primary tool/vendor documentation (not a blog or listicle)
- `news` — recent news/benchmark article; informs tool discovery only, never the sole ground for a factual claim
---
## References
- [1] Consensus (2025). *Welcome to Consensus*. https://consensus.app/home/blog/welcome-to-consensus/ — accessed 2026-07-27. [type: primary-doc] — Grounds the semantic-vs-keyword slide from a working system's own description: retrieval uses "a hybrid search approach that combines" "Semantic search, powered by AI embeddings, to capture the intent behind your question" and "Keyword search, powered by BM25, a proven traditional method, [which] anchors results to the exact terms in your query for precise keyword matching"; then re-ranks "the top 1,500 papers" on "recency of publication, citation count, journal reputation and impact" before a final precision pass over "the top 20 papers". Also grounds the corpus figure ("over 200 million scientific documents", aggregated from Semantic Scholar, OpenAlex and "our own crawl of the scholarly web"), the design rule "We only use AI after we search the scientific literature", and the three-way hallucination taxonomy (fake sources / wrong facts / misread sources) with the vendor's claim that in a retrieval-first architecture "only the third type of hallucination is possible". **Currency note (2026-09-02 revision):** this is the vendor's 2025 description; Consensus has since rebuilt around a GPT-5-based multi-agent "Scholar Agent" [35], so the pipeline is taught in the past tense as a worked example of retrieve-then-generate, not as the current architecture.
- [2] Bramer, W. M., Giustini, D., & Kramer, B. M. R. (2016). *Comparing the coverage, recall, and precision of searches for 120 systematic reviews in Embase, MEDLINE, and Google Scholar: a prospective study*. Systematic Reviews 5, 39. https://pmc.ncbi.nlm.nih.gov/articles/PMC4772334/ — accessed 2026-07-27. [type: peer-reviewed] — Grounds "coverage is not recall" and the Google Scholar caution: across 4,795 included references from 120 systematic reviews, Google Scholar's coverage was 97.2% (Embase/MEDLINE combined 97.5%, MEDLINE alone 92.3%), but recall for Embase/MEDLINE combined was 81.6% versus 72.8% for Google Scholar, and only 46.4% of included references were among Google Scholar's downloadable first 1,000 results. Precision in Google Scholar's first 1,000 was 1.9%. Conclusion quoted on the caution-zone slide: "neither GS nor one of the other databases investigated, is on its own, an acceptable database to support systematic review searching."
- [3] Lau, O., & Golder, S. (2025). *Comparison of Elicit AI and Traditional Literature Searching in Evidence Syntheses Using Four Case Studies*. Cochrane Evidence Synthesis and Methods 3(5), e70050. https://doi.org/10.1002/cesm.70050 — accessed 2026-07-27. [type: peer-reviewed] — The load-bearing independent evaluation in this session. Using "the subscription-based version of Elicit Pro in Review mode" against four completed evidence syntheses (evaluations run 24 Feb – 11 Mar 2025), Elicit's sensitivity averaged 39.5% (range 25.5–69.2%) against 94.5% (91.1–98.0%) for the original searches, while its precision averaged 41.8% (35.6–46.2%) against 7.55% (0.65–14.7%). Crucially for the "each finds what the other misses" claim, Elicit "identified some included studies not identified by the original searches". Authors' conclusion: Elicit "is not sensitive enough to replace traditional searching" but its high precision makes it useful for preliminary searches and as an adjunct. **Currency note (2026-09-02 revision):** these figures are a **Feb–Mar 2025 measurement of Elicit Pro's Review mode** and are labelled as such wherever they appear; Elicit has since shipped a substantially different product (Research Agent, 2026-08-04 [33]), and no independent evaluation of that product was found as of 2026-09-02. The deck's earlier "the one independent evaluation of any tool on this list" uniqueness phrasing was retired in this revision: Elicit now publishes its own agent benchmarking (vendor-run and self-graded, so excluded from slides under the standing rule below), and [31] independently audits the deep-research tier.
- [4] Sahu, G., Charlin, L., & Pal, C. (2026). *Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth*. arXiv:2605.29234. https://arxiv.org/abs/2605.29234 — accessed 2026-07-27. [type: peer-reviewed] — Grounds two slides. First, that agentic breadth-first expansion along bibliographies "substantially outperforms vanilla API-only search, raising recall on ROLLINGEVAL-JUN25 (a 250-paper literature-search benchmark) from below 20% to above 80%". Second, the humbling counterpoint used on the same slide: "only 51% of human citations are judged moderately relevant or higher, against 86–88% for the strongest AI-based re-rankers", and on the OpenAlex co-authorship graph "humans are 2.5× more likely than the best AI re-rankers to cite a direct collaborator". Verified via Exa (arXiv, 2026-05-28, Mila/HEC Montréal/ServiceNow Research).
- [5] Pan, T., Cheng, M., Wang, D., Zhou, Y., Ouyang, J., Liu, Q., & Chen, E. (2026). *ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments*. arXiv:2606.20235. https://arxiv.org/abs/2606.20235 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the ceiling on agentic search: built from "over 1,000 computer science topics and four representative research intents", "agentic methods outperform single-shot retrieval baselines, yet the best-performing agent only achieves 0.314 Recall@100 and 0.355 Recall@All, indicating substantial room for improvement". Used on the agent-limitations slide as the sober counterweight to [4]'s optimistic recall figure. Verified via Exa (arXiv, 2026-06-18, USTC).
- [6] Elicit (2026). *Elicit: AI for scientific research*. https://elicit.com/ — accessed 2026-07-27. [type: primary-doc] — Primary description of Elicit for the comparison matrix: search, summarize, extract data from and chat with "over 125 million papers"; research reports "based on a process inspired by systematic reviews"; "Elicit supports all AI-generated claims with sentence-level citations from the underlying sources"; "Elicit can find up to 1,000 relevant papers and analyze up to 20,000 data points at once". Note the internal inconsistency used on the "read the corpus number critically" slide: this page says 125 million while the vendor's own pricing page [7] says "more than 138 million".
- [7] Elicit (2026). *Pricing | Elicit*. https://elicit.com/pricing — accessed 2026-07-27. [type: primary-doc] — Grounds Elicit's access/cost cell: Basic tier free with "limited usage for Research Agent and Research Reports" and "unlimited search across more than 138 million papers"; Pro $49 per user/month billed as $588 annually, with a systematic-review workflow that "can screen 5,000 papers"; Scale $169 per user/month; Enterprise custom. Also the source of the 138M corpus figure that conflicts with [6].
- [8] Elicit (2026). *Export your data from Elicit*. https://support.elicit.com/en/articles/1153857 — accessed 2026-07-27. [type: primary-doc] — Grounds Elicit's citation-export cell: "Results from Find Papers, Extract Data, Paper Chat, and Research Reports can be exported as .bib or .ris files"; reports export as PDF or Word on any plan; table export requires Elicit Plus or higher; and the limitation stated verbatim on the handout: "Elicit's library cannot be exported at this time."
- [9] Consensus (2026). *Pricing — Consensus*. https://consensus.app/home/pricing/ — accessed 2026-07-27. [type: primary-doc] — Grounds Consensus's access/cost cell: Free tier with "unlimited searches across over 200M research papers" plus 10 Pro Analyses, 10 Study Snapshots and 10 Ask Paper messages per month; Premium $11.99/month ($108 billed annually); Teams $12.99 per seat/month; Enterprise custom; 40% student discount on a verified .edu or .ac address.
- [10] Consensus (2026). *How to Export Consensus Results to Reference Managers (EndNote, Zotero, Paperpile)*. https://help.consensus.app/en/articles/9922811-how-to-export-consensus-results-to-reference-managers-endnote-zotero-paperpile — accessed 2026-07-27. [type: primary-doc] — Grounds Consensus's citation-export cell: "Export paper details from search results and saved lists in Consensus to CSV or RIS file formats. These formats are compatible with popular reference managers like EndNote, Mendeley, Zotero, and more."
- [11] Semantic Scholar / Ai2 (2026). *About Semantic Scholar*. https://www.semanticscholar.org/about — accessed 2026-07-27. [type: primary-doc] — Grounds the Semantic Scholar row: "We index over 200 million academic papers sourced from publisher partnerships, data providers, and web crawls"; operated by Ai2, "a non-profit research institute founded in 2014"; publishes the Semantic Scholar Academic Graph (S2AG) dataset and the Semantic Scholar Open Research Corpus (S2ORC), "the largest publicly-available collection of machine-readable academic text".
- [12] Semantic Scholar (2026). *Semantic Scholar Academic Graph API — Overview*. https://www.semanticscholar.org/product/api — accessed 2026-07-27. [type: primary-doc] — Grounds the precise corpus figures and the "free scholarly infrastructure" point: "214 Million Papers", "2.49 Billion Citations", "79 Million Authors"; the API exposes "authors, papers, citations, venues, SPECTER2 embeddings"; "Most Semantic Scholar endpoints are available to the public without authentication". Also the source of the observation that several other tools in the matrix are built on this data.
- [13] Scite (2026). *AI for Research | Scite*. https://scite.ai/ — accessed 2026-07-27. [type: primary-doc] — Grounds the Scite row: "Reads across 280M+ full-text scholarly articles"; "Search across 280M+ articles, preprints, books, patents, and datasets"; Smart Citations "show how later research has supported, challenged, or discussed a paper, claim, method, journal, author, institution, or funder"; "built on direct agreements with Wiley, SAGE, and 30+ more publishers" and elsewhere on the same page "agreements across 40+ publishers plus the full open access corpus"; "Every claim Scite's AI makes links back to the specific sentence in the specific paper it came from"; runs "inside Claude, ChatGPT, and any MCP-compatible assistant, plugs into Zotero, and is available through our API".
- [14] Scite (2026). *Scite Pricing*. https://scite.ai/pricing — accessed 2026-07-27. [type: primary-doc] — Grounds Scite's access/cost cell (Basic $20/month, Pro $50/month, Team $50/seat/month, Enterprise custom, 7-day free trial; no free standing tier) and, on the "read the corpus number critically" slide, the internal inconsistency with [13]: this page states "300M+ Scholarly sources" where the homepage states 280M+.
- [15] ResearchRabbit (2026). *ResearchRabbit*. https://www.researchrabbit.ai/ — accessed 2026-07-27. [type: primary-doc] — Grounds the ResearchRabbit row: "Access over 310 million academic papers"; the citation-graph proposition, "See how topics relate to each other and evolve over time by using built-in visualizations"; and seed-paper expansion, "Start with one paper and easily expand your search by adding new authors, related works, and emerging topics."
- [16] ResearchRabbit (2026). *Pricing | ResearchRabbit*. https://www.researchrabbit.ai/pricing — accessed 2026-07-27. [type: primary-doc] — Grounds ResearchRabbit's access/cost cell: "Free — $0, Forever!" with "unlimited searches across 310+ million articles", unlimited library and collections, and "up to 50 seed articles"; ResearchRabbit+ at $10/month annual or $12.50/month monthly raising the seed limit to 300, with country-parity discounts; an Institution tier with LibKey integration.
- [17] ResearchRabbit (2026). *Using the Zotero Importer*. https://learn.researchrabbit.ai/en/articles/12796541-using-the-zotero-importer — accessed 2026-07-27. [type: primary-doc] — Grounds the honest limitation in ResearchRabbit's citation-export cell: the Zotero integration is currently one-directional import, with export back out via BibTeX only — "you can bring articles you find in ResearchRabbit back into Zotero using the Export feature – Zotero will happily import the BibTeX file format" — and "We're actively working on a two-way sync feature", i.e. two-way sync is announced but not yet shipped.
- [18] SCiNiTO (2026). *What is SCiNiTO?* https://www.scinito.ai/help/getting-started/what-is-scinito — accessed 2026-07-27. [type: primary-doc] — Grounds the SCiNiTO row, included as one option among peers: "an AI research assistant built on top of OpenAlex, the open scholarly catalog", bundling "search, an AI chat that cites real sources, PDF analysis, a journal recommender, an AI reviewer for manuscripts, and shared research projects ('Spaces')"; "Search across 500M+ academic works with filters for date, journal, open access, SDGs, and more"; plans are Guest (5 search queries and 5 AI Chat prompts), Individual free ("unlimited smart search, enhanced AI, bookmarks, Spaces") and Institutional. The 500M+ figure is vendor-stated and is one of the numbers interrogated on the "read the corpus number critically" slide against OpenAlex's own published count [20].
- [19] SCiNiTO (2026). *Boolean operators (AND, OR, NOT)*. https://www.scinito.ai/help/smart-search/boolean-operators — accessed 2026-07-27. [type: primary-doc] — Grounds the "some tools give you both modes" point on the semantic-vs-keyword slide: "SCiNiTO supports boolean operators inside the Smart Search bar", with nesting (`(public health OR healthcare) AND AI NOT survey`), quoted phrases for exact matching, and Boolean queries combinable with filters such as open access.
- [20] OpenAlex (2026). *About OpenAlex*. https://openalex.org/about — accessed 2026-07-27. [type: primary-doc] — Grounds the paywall caution zone with a first-party number: OpenAlex's own comparison table reports 243M works of which 48M are open access (roughly one in five), against Scopus at 87M works / 20.5M OA and Web of Science core at 87M / 12M. Also grounds the "open infrastructure" point — CC0, non-profit, freemium, "our data is free and reusable, available via bulk download or API" — and provides the published catalogue size against which SCiNiTO's 500M+ claim [18] is checked.
- [21] OpenAI (2026). *Deep research in ChatGPT*. https://help.openai.com/en/articles/10500283-deep-research-in-chatgpt — accessed 2026-07-27. [type: primary-doc] — Grounds what a deep-research agent does, in the vendor's own words: "ChatGPT creates a proposed research plan. You can review and modify it before the research begins"; the user chooses sources ("the public web", uploaded files, connected apps) and can restrict research to named sites; "You get a structured report with citations or source links so you can verify the information"; and the vendor's own scoping advice, "Use search for quick facts, and use deep research for depth and thoroughness." **Currency note (2026-09-02 revision):** OpenAI's published per-tier deep-research quotas trace to mid-2025 documentation and the Pro plan was restructured in April 2026, so no quota figure is stated as current anywhere in this session; per secondary reporting [34], deep research has run a GPT-5.2-based model (not the GPT-5.6 chat flagship) since February 2026.
- [22] Google (2026). *Gemini Deep Research Agent — Gemini API*. https://ai.google.dev/gemini-api/docs/interactions/deep-research — accessed 2026-07-27. [type: primary-doc] — Grounds the agent-mechanics slide: an agent that "autonomously plans, executes, and synthesizes multi-step research tasks" and "navigates complex information landscapes to produce detailed, cited reports"; "Research tasks involve iterative searching and reading and can take several minutes to complete"; collaborative planning "gives you control over the research direction before the agent starts its work by letting you review and refine the research plan before execution". Still labelled Preview as of the 2026-07-27 fetch (not re-fetched since); the underlying model and the Deep Research Max variant are grounded separately in [32].
- [23] Anthropic (2026). *Using Research on Claude*. https://support.claude.com/en/articles/11088861-using-research-on-claude-ai — accessed 2026-07-27. [type: primary-doc] — Third vendor's primary description of the same category, so the deck does not characterise deep research from one vendor alone: "Claude operates agentically, conducting multiple searches that build on each other while determining exactly what to investigate next"; it "delivers thorough answers in minutes, complete with easy-to-check citations"; available on paid plans only (Pro, Max, Team, Enterprise) and requires web search to be enabled.
- [24] Venkit, P. N., Laban, P., Zhou, Y., Huang, K.-H., Mao, Y., & Wu, C.-S. (2025). *DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence*. arXiv:2509.04499. https://arxiv.org/abs/2509.04499 — accessed 2026-07-27. [type: peer-reviewed] — The primary audit source for the documented-failure-modes slide. An audit framework of "eight measurable dimensions spanning answer text, sources, and citations" applied to public systems including GPT-4.5/5, You.com, Perplexity, Copilot/Bing and Gemini: systems "frequently produce one-sided, highly confident responses on debate queries and include large fractions of statements unsupported by their own listed sources", with citation accuracy ranging "40–80% across systems"; deep-research configurations reduce overconfidence and improve citation thoroughness but "remain highly one-sided on debate queries and exhibit large fractions of unsupported statements". **Currency note (2026-09-02 revision):** kept as the **2025 baseline audit** — the audited-systems list (a quotation from the paper) names GPT-4.5/5-era deployments, and the paper has a single version (v1, 2025-09-02) with no update as of 2026-09-02. The current-generation companion audit is [31]; the one-sidedness finding has **no** current-generation re-measurement either way, and is therefore stated in this session as a 2025 finding that nothing newer has overturned.
- [25] Rao, D., Wong, E. W. M., & Callison-Burch, C. (2026). *Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents*. arXiv:2604.03173. https://arxiv.org/abs/2604.03173 — accessed 2026-09-02 (PDF re-fetched; first accessed 2026-07-27). [type: peer-reviewed] — Carried forward from Session 1 (where it is source [8]) so the series states this figure consistently. Across 10 models and agents on DRBench (53,090 URLs) and 3 models on ExpertQA (168,021 URLs across 32 academic fields) — 221k+ URLs in total: "3–13% of citation URLs are hallucinated — they have no record in the Wayback Machine and likely never existed — while 5–18% are non-resolving overall"; "Deep research agents generate substantially more citations per query than search-augmented LLMs but hallucinate URLs at higher rates"; from the body, verified verbatim: "Pooling across the two deep research agents, the hallucination rate is 10.7% [10.2, 11.2] versus 4.8% [4.3, 5.2] for the eight search-augmented models" (z = 15.15, p < 10⁻⁵¹) — the number behind the "depth amplifies URL fabrication" point. Non-resolving rates range "from 5.4% (Business) to 11.4% (Theology)". Also grounds the optimistic half of the verification slide: automated URL liveness checking reduced non-resolving citation URLs "by 6–79× to under 1%".
- [26] Walters, W. H., & Wilder, E. I. (2023). *Fabrication and errors in the bibliographic citations generated by ChatGPT*. Scientific Reports 13, 14045. https://www.nature.com/articles/s41598-023-41032-5 — accessed 2026-07-27. [type: peer-reviewed] — Carried forward from Session 1 (source [5] there) to keep the fabrication baseline consistent across the series: across 636 citations in 84 generated literature reviews, "55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated", and "43% of the real (non-fabricated) GPT-3.5 citations but just 24% of the real GPT-4 citations include substantive citation errors". Used on this session's slides for the second half of that finding — a citation can be real and still be wrong about what it says — always as a **labelled 2023 baseline (GPT-3.5 / GPT-4 era)**, never as a current rate; the current-generation counterpart of "real but wrong about what it says" is [31].
- [27] Flemyng, E., Noel-Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., Jordan, Z., Minx, J., Eisele-Metzger, A., Hamel, C., Jemioło, P., Porritt, K., & Grainger, M. (2025). *Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence 2025*. Cochrane Database of Systematic Reviews, Editorial ED000178. https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.ED000178/full — accessed 2026-07-27. [type: policy] — Grounds the systematic-review caution zone with the field's own governing statement rather than opinion. Key messages quoted on the slide: "Evidence synthesists are ultimately responsible for their evidence synthesis, including the decision to use artificial intelligence (AI) and automation"; authors "can use AI and automation as long as they can demonstrate that it will not compromise the methodological rigour or integrity of their synthesis"; "AI and automation in evidence synthesis should be used with human oversight"; and "Any use of AI or automation that makes or suggests judgements should be fully and transparently reported in the evidence synthesis report." The statement endorses the RAISE (Responsible use of AI in evidence SynthEsis) recommendations. Co-published across the four organisations' journals.
- [28] Lieberum, J.-L., Toews, M., Metzendorf, M.-I., Heilmeyer, F., Siemens, W., Haverkamp, C., Böhringer, D., Meerpohl, J. J., & Eisele-Metzger, A. (2025). *Large language models for conducting systematic reviews: on the rise, but not yet ready for use — a scoping review*. Journal of Clinical Epidemiology 181, 111746. https://www.jclinepi.com/article/S0895-4356(25)00079-4/fulltext — accessed 2026-07-27. [type: peer-reviewed] — Carried forward from Session 1 (source [17] there), where it grounds the "Unreliable" zone of the Capability/Failure Map. Here it grounds the empirical picture behind the systematic-review caution zone: 37 included articles covering 10 of 13 defined systematic-review steps, most often literature search (41%), study selection (38%) and data extraction (30%); "about half of the studies found LLMs promising, a quarter were neutral, and one-fifth found them not promising"; conclusion that "LLMs should only be used with caution and under human supervision".
- [29] Maddi, A., Maisonobe, M., & Boukacem-Zeghmouri, C. (2025). *Geographical and disciplinary coverage of open access journals: OpenAlex, Scopus, and WoS*. PLOS ONE 20(3), e0320347. https://doi.org/10.1371/journal.pone.0320347 — accessed 2026-07-27. [type: peer-reviewed] — The primary grounding for the coverage-bias caution zone. Against a ROAD baseline of 62,701 active open-access journals (May 2024), Web of Science indexes 6,157, Scopus 7,351 and OpenAlex 34,217, with 24,976 journals exclusive to OpenAlex and only 4,094 indexed in all three. Continental coverage indices (neutral = 1) show Oceania at 4.30 in WoS and 3.39 in Scopus against 1.46 in OpenAlex, Africa consistently underrepresented across all three, and "consistent over-representation" of high-income economies in Scopus and WoS while low- and middle-income countries remain "particularly underrepresented".
- [30] Culbert, J., Hobert, A., Jahn, N., Haupka, N., Schmidt, M., Donner, P., & Mayr, P. (2025). *Reference coverage analysis of OpenAlex compared to Web of Science and Scopus*. Scientometrics. https://doi.org/10.1007/s11192-025-05293-3 (preprint: https://arxiv.org/abs/2401.16359) — accessed 2026-07-27. [type: peer-reviewed] — The counterweight that keeps the coverage-bias slide fair rather than alarmist: on a cleaned dataset of 16.8 million recent publications shared by all three databases, OpenAlex has "average source reference numbers and internal coverage rates comparable to both Web of Science and Scopus". OpenAlex captures more ORCID identifiers, fewer abstracts, and a similar number of open-access status indicators per article. Grounds the slide's conclusion that the open catalogue is not simply worse — it is differently shaped.
- [31] Onweller, H., Lumer, E., Huber, A., Ramchandani, P., Subbiah, V. K., & Feld, C. (2026). *Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents*. **Preprint** (arXiv:2605.06635, v1 submitted 2026-05-07). https://arxiv.org/abs/2605.06635 — accessed 2026-09-02. [type: peer-reviewed] — Labelled a preprint wherever cited. The **current-generation companion to the DeepTRACE 2025 baseline [24]**, added in the 2026-09-02 revision: 14 closed- and open-source models (including Claude Opus 4.5/4.6, GPT-5.2/5.4, Gemini 3.1 Pro) audited on three dimensions — link accessibility, topical relevance, factual accuracy against the retrieved source. Quoted verbatim from the abstract: "even the strongest frontier models maintain link validity above 94% and relevance above 80%, yet achieve only 39-77% factual accuracy"; and on research depth, "Fact Check accuracy drops by approximately 42% on average across two frontier models as tool calls scale from 2 to 150, demonstrating that more retrieval does not produce more accurate citations." Grounds the demotion of the resolve check from sufficient to necessary on slide 26 and checklist item 1: the dominant failure has moved from *fake URL* to *real URL, unsupported claim*. Session 6 cites the same paper (as its [66]) with the same figures.
- [32] Google (2026). *Next-generation Gemini Deep Research* (Google blog, 2026-04-21). https://blog.google/innovation-and-ai/models-and-research/gemini-models/next-generation-gemini-deep-research/ — accessed 2026-09-02. [type: primary-doc] — Grounds the current identity of Google's Tier-3 agents: "Built with Gemini 3.1 Pro, the new Deep Research agents bring MCP support, native visualizations and unprecedented analytical quality to long-horizon research workflows"; two variants — Deep Research (speed/efficiency) and Deep Research Max, which "is focused on long-term research tasks with high accuracy requirements" using additional test-time computation. The launch post's DeepSearchQA and HLE score claims are vendor self-reported and therefore appear nowhere in this session, per the standing rule.
- [33] Elicit (2026). *Elicit Blog* (release index: "Introducing Elicit Research Agent", 2026-08-04; "The Elicit API and MCP", 2026-07-15; usage-pool change, 2026-06-30). https://elicit.com/blog — accessed 2026-09-02. [type: primary-doc] — Grounds two facts that changed the currency of [3] and the tier-boundary caveat: Elicit's product is no longer the early-2025 Review mode that Lau & Golder measured — "Elicit Research Agent" launched 2026-08-04 as "an AI environment built for high-stakes decisions" producing "fully cited reports, data analyses, tables, visualizations, slides, and documents" — and Elicit ships an API + MCP server (2026-07-15), so Tier-1 chatbots can call it as a tool. The launch post also reports vendor-run benchmarking of the agent against a frontier chatbot baseline; that figure is vendor self-reported and self-graded, and is excluded from slides, handout and demo under the standing rule below — its existence is cited only as the reason the deck's former "one independent evaluation" uniqueness claim was retired.
- [34] Wikipedia (2026). *Deep research* (article, citing OpenAI documentation). https://en.wikipedia.org/wiki/Deep_research — accessed 2026-09-02. [type: news] — Secondary source, used for one dated product fact and labelled as secondary where cited: "Since February 2026, it [ChatGPT deep research] uses a model based on OpenAI's GPT-5.2, having originally been released with a specialized version of OpenAI's o3." Teaching point on slide 17: the deep-research button and the headline chat model (GPT-5.6) are not the same model. The same article's per-tier quota figures trace to June-2025 OpenAI documentation and are treated as stale (see [21] currency note); OpenAI's own help pages refused automated fetch on 2026-09-02, which is why a secondary source carries this fact — flagged rather than hidden.
- [35] OpenAI (2026). *Consensus: building Scholar Agent on GPT-5* (customer case study). https://openai.com/index/consensus/ — accessed 2026-09-02. [type: primary-doc] — Grounds the tier-boundary caveat's Consensus half: Consensus rebuilt its product around "Scholar Agent... a multi-agent system built on GPT-5 and the Responses API", with distinct planning, search, reading and analysis agents. Cited for the descriptive architecture fact only — it is an OpenAI marketing case study, and no performance claim from it appears anywhere in this session. Read alongside [1]: the 2025 pipeline description this deck teaches from is now a dated description of a rebuilt system, and slide 7 says so.
- [36] Consensus (2026). *Consensus MCP Server* (documentation). https://docs.consensus.app/docs/mcp — accessed 2026-09-02. [type: primary-doc] — Grounds the Consensus leg of the tier-boundary caveat: the MCP server lets ChatGPT, Claude and other MCP clients "search over 200 million peer-reviewed academic research papers directly from the conversation". Together with Scite's MCP [13] (shipped 2026-02-26) and Elicit's API + MCP [33] (2026-07-15), this is why the taxonomy's Tier 2 / Tier 3 boundary is taught as blurring: a Tier-1 chatbot can now call a Tier-2 scholarly corpus mid-conversation. The taxonomy survives as a map of what you are reaching for, not of which product you opened.
<!--
Counts for this session (re-checked 2026-09-02 after the August-2026 modernization pass):
total entries = 36 (target >= 15)
peer-reviewed/preprint = 11 ([2] [3] [4] [5] [24] [25] [26] [28] [29] [30] [31]; target >= 5)
policy = 1 ([27])
primary-doc = 23 ([1] [6] [7] [8] [9] [10] [11] [12] [13] [14] [15] [16] [17] [18] [19] [20] [21] [22] [23] [32] [33] [35] [36])
news = 1 ([34] — secondary reporting of OpenAI documentation; flagged as secondary where cited)
Sources [25], [26] and [28] are deliberately shared with Session 1 so the series
quotes the same numbers for the same claims across sessions; [31] is shared with
Session 6 (its [66]) for the same reason.
Numbering was compacted 1-30 in the fix pass when the former entry [9] was
removed by review ruling; former [10]-[31] are now [9]-[30]. Entries [31]-[36]
were appended in the 2026-09-02 modernization pass; no existing number changed.
Accessed dates: entries re-fetched during the 2026-09-02 research pass carry
that date; all others keep their original 2026-07-27 date, per the series
accessed-date-honesty rule.
-->
### Review ruling (Ehsan checkpoint, 2026-07-27): no vendor self-reported performance figures
An earlier draft of this session carried a 31st entry — Elicit's own blog post *Evaluating Elicit's Systematic Literature Review Capabilities* — used on slide 20 as evidence of what the vendor claims about itself and how it measured (95.0% search recall across 994 Cochrane reviews, using the review title as the whole query and counting a review's excluded studies as relevant).
**Ruling: keep the lesson, drop the vendor figure.** The vendor blog post has been removed from this list entirely, and the numbering compacted so the sequence runs 1–30 with no gaps.
The lesson survives, rebuilt on independent evidence only. Slide 20 now makes the same point — *a headline performance number is meaningless without its methodology* — from the spread **within** the independent peer-reviewed evaluation [3]: the same tool, in the same mode, evaluated by the same team over about two weeks, scored between 25.5% and 69.2% sensitivity depending on which review it was tested against, so the widely quoted 39.5% is a mean of four case studies that describes none of them. The audited 40–80% citation-accuracy band across deep-research systems [24] makes the same point one tier up. Both figures are peer-reviewed/preprint sources whose wording was confirmed at the primary source.
**Standing rule for this deck:** no vendor's self-reported performance claim appears on any slide, in the handout, or in the demo script. Vendor documentation is cited only for descriptive facts about a tool — corpus size, features, export formats, pricing — and every such figure is labelled on-slide as vendor-stated.
### Dropped during research (and why)
- *"Assessing the Effectiveness of AI Tools (Elicit, SciSpace, and Consensus) in Literature Review and Research"*, Canadian Journal of Information and Library Science — the single most on-topic-looking hit in the whole search, and **retracted**. Dropped entirely. (Mentioned as a live example in `demo-script.md`'s presenter notes, because "the top search result was a retracted paper" is exactly the failure this session teaches.)
- Comparison listicles and affiliate-style tool roundups (`iatrox.com`, `paperguide.ai`, `thedrive.ai`, `sourcely.net`, `atlasworkspace.ai`, `perplexityaimagazine.com`, `agentaya.com`) — several carried plausible-looking corpus and feature figures that disagreed with the vendors' own pages. Every cell in the comparison matrix is taken from the tool's own documentation instead.
- Third-party corpus figures such as "Elicit's 138M-paper Semantic Scholar database" — this conflates Elicit's corpus with Semantic Scholar's and is contradicted by both vendors' own pages ([6], [7], [12]). Dropped; the matrix cites each vendor for its own number and flags where a vendor contradicts itself.
- General "semantic search beats keyword search on precision and recall" claims from computer-science comparison papers on general-purpose web search engines — not transferable to scholarly retrieval, and contradicted for this use case by [3]. The semantic-vs-keyword slides are grounded in scholarly-retrieval evidence only ([1], [2], [3], [4]).
- The Scientometrics paper on regional disparities in WoS and Scopus journal coverage (10.1007/s11192-024-04948-x) — could not be fetched past the publisher's authentication redirect, so its figures were not confirmed at source. The same claim is made instead from [29], which is open access and whose figures were read directly.
- A widely repeated "Elicit searches 126 million papers and uses GPT-3" description — traceable only to older secondary write-ups. Dropped in favour of the vendor's current pages.
- OpenAI's *Introducing deep research* announcement page — would not load through either fetcher on 2026-07-27, so nothing from it is cited. The category description rests on the three vendors' current product documentation ([21], [22], [23]) instead.
---
*AI for Researchers · Session 2: Literature Discovery & Synthesis · Landscape as of August 2026*