← series index
# Sources — Session 3: Reading, Notes & Knowledge Management
*AI for Researchers · dated, annotated reference list · compiled 2026-07-27, revised 2026-08-21 · Landscape as of August 2026*
Entry format (one line per source, numbering matches the `[n]` footnote markers used in `slides.html` and `handout.md`):
`- [n] Author/Org (Year). *Title*. URL — accessed YYYY-MM-DD. [type: peer-reviewed | policy | primary-doc | news] — one-line note on what claim(s) it grounds.`
Allowed `type` values:
- `peer-reviewed` — peer-reviewed paper or arXiv/preprint literature
- `policy` — official publisher/funder/institutional policy page
- `primary-doc` — primary tool/vendor documentation (not a blog or listicle)
- `news` — recent news/benchmark article; informs tool discovery only, never the sole ground for a factual claim
---
## References
### A. The critical-literacy evidence base (where summaries flatten nuance)
- [1] Peters, U., & Chin-Yee, B. (2025). *Generalization bias in large language model summarization of scientific research*. Royal Society Open Science 12(4), 241776. https://doi.org/10.1098/rsos.241776 — accessed 2026-08-21. [type: peer-reviewed] — **The load-bearing study of the critical-literacy segment.** Across 4,900 LLM-generated summaries (4,300 of abstracts, 600 of full articles) from 10 prominent LLMs including ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B and Claude 3.7 Sonnet: LLM summaries were "twice as likely to contain generalized conclusions compared to the original abstracts"; DeepSeek, ChatGPT-4o and LLaMA 3.3 70B overgeneralized in "26–73% of cases"; prompting explicitly for accuracy made overgeneralization *more* likely, not less (odds ratio 1.90); and against expert-written NEJM Journal Watch summaries LLM summaries were nearly five times more likely to contain broad generalizations (OR = 4.85, 95% CI [3.06, 7.70], p < 0.001). Also grounds the counter-intuitive line on the slide that "newer models tended to perform worse in generalization accuracy than earlier ones" — **now framed on the slide as a within-cohort observation about the 2025 model generation, not a law**: the deck pairs it with the same team's 2026 follow-up [33], and no study yet tests the August-2026 cohort on overgeneralization. Also grounds the mechanism — the most frequent shift was converting quantified generalizations into generics (from "the patients in this trial improved" to "the treatment improves patients").
- [33] Peters, U., Bertazzoli, A., DeJesus, J. M., van der Velden, G. J., & Chin-Yee, B. (2026). *Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models*. Public Understanding of Science, first published online 20 April 2026. https://doi.org/10.1177/09636625261425891 — accessed 2026-08-21. [type: peer-reviewed] — **The same team's direct follow-up to [1], one model generation later** — added in the August-2026 revision so the deck does not rest its central claim on a 2025 cohort alone. Comparing laypeople, scientists and two LLMs on the interpretation of generics ("statins reduce cardiovascular events"): ChatGPT-5 rated generics as *more* generalizable than laypeople (b = .25, p = .009) and DeepSeek-V3.1 more so (b = .53, p < .001), with inflated credibility ratings (b = 1.05 and b = .69 respectively, both p < .001) — while human domain experts went the other way, rating generics as *less* generalizable than laypeople (psychologists b = −.32; biomedical researchers b = −.48; both p < .001). Grounds the slide-10 line that overgeneralization persists on GPT-5-class models and is directional: experts read a generic claim narrowly, LLMs read it more broadly than even untrained laypeople. Caveat carried on the slide: still one generation behind the August-2026 flagships — that gap is stated as open, not resolved.
- [2] Hammond, K., & Newell, S. (2025). *Reading between the GenAI lines: Twenty authors' perceptions of GenAI-summaries of their research and the implications for learning and teaching*. ASCILITE Publications (refereed conference poster, ASCILITE 2025, Adelaide), pp. 25–26. https://doi.org/10.65106/apubs.2025.2719 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the "what exactly gets dropped" slide with evidence from the only people who can adjudicate it: the authors of the papers being summarised. Semi-structured interviews with 20 Australasian-resident journal-article authors, comparing their own summary of their paper with a ChatGPT-4 summary of it, analysed by Reflexive Thematic Analysis. Findings quoted on the slide: summaries were "accurate but vague"; authors identified "inaccuracies and omissions of: context, disciplinary conventions and concepts, important findings and limitations, and connection/attribution of ideas to prior authors"; "AI-summaries present author suggestions as established facts"; "The absence of methodological details was concerning because methods sections provide information on researchers' decisions, validity of the findings, and training for novice researchers/students"; in introduction summaries "prior literature was omitted, thus misrepresenting the study as an isolated unit instead of ongoing knowledge-building conversations". Note on weight: this is a two-page refereed poster abstract reporting a qualitative study of 20 authors — it is cited on the slide as qualitative, author-side evidence about *which* content disappears, never as a frequency estimate.
- [3] Etkin, H. K., Etkin, K. J., Carter, R. J., & Rolle, C. E. (2025). *Differential effects of GPT-based tools on comprehension of standardized passages*. Frontiers in Education 10, 1506752. https://doi.org/10.3389/feduc.2025.1506752 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the "what it costs the reader" slide. Pre-registered randomised cross-over online study, n = 195 college-aged participants, ACT-derived reading passages, four GPT-based tools (AI-generated summaries, AI-generated outlines, a Q&A tutor chatbot, a Socratic discussion chatbot). Result quoted verbatim on the slide: "AI tools significantly improved comprehension in lower performing participants and significantly worsened comprehension in higher performing participants." The summary tool was the *worst* tool for strong readers (t(123) = 9.2, p < 0.001, d = 0.83) and the weakest of the four for low performers (d = 0.45). The authors' own explanation is the sentence the segment is built around: reading a summary instead of the passage worsened comprehension in higher performers "likely because much of the detail and nuance of the passage was lost in the summary". Caveat stated on the slide: standardised test passages, not journal articles — the effect direction transfers, the effect size should not be quoted for research reading.
- [4] Thelwall, M. (2026). *Will AI be overconfident about academic research findings when reliant on abstracts?* arXiv:2605.27392. https://arxiv.org/abs/2605.27392 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the "abstract trap" slide, which pairs with the primary-doc fact [19] that most extraction tools fall back to the abstract when they cannot reach the PDF. Full-text articles were submitted to GPT-OSS 120B, which rated the strength of the claim for the main result separately in the abstract, the discussion and the conclusion. Finding quoted verbatim: "Outside the social sciences and humanities, claims tended to be stronger in the abstract and conclusions than the discussion, suggesting that relying on the strength of claims in abstracts would be misleading." The author's framing is the one the slide uses: this is "a weak hallucination: not stating a false fact but misrepresenting the strength of evidence for it". Preprint, single-model design — cited for direction, not for a magnitude.
- [5] Martin, A., Humphreys, W., Brown, M. C., Leckey, C. A. C., & Kaur, H. (2026). *An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems*. arXiv:2602.21059. https://arxiv.org/abs/2602.21059 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the "what experts actually catch" slide and the handout's triage checklist. Two-phase qualitative study: 20 error patterns across seven categories identified from thematic analysis of 68 question-answer pairs with domain experts (N=3) evaluating a scholarly QA system on papers they had themselves authored, then validated through contextual inquiries with 10 further scientists across science and engineering domains. The two findings used on the slide: the error taxonomy runs from fabricated citations and invented technical terms through to "synthesis failures in multi-document contexts"; and "while experts naturally identified errors in correctness and completeness, the structured schema appeared to help them detect previously overlooked issues, particularly subtle hallucinations and citation errors" — i.e. a checklist measurably out-performs unaided expert reading, which is the justification for issuing one.
- [10] Xu, M., Li, X., Lan, T., Zeng, X., Wang, T., Sheng, S., et al. (2026). *Comparing structural and conversational AI summarisation for time-constrained academic reading support among graduate students*. Interactive Learning Environments. https://doi.org/10.1080/10494820.2025.2604649 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the honest caveat on the demo-preview slide: a within-subject user study (n = 29 HCI graduate students) comparing conversational AI summarisation, structural AI summarisation and no AI on 10-minute paper-categorisation tasks found that "participants in AI-assisted conditions reported higher user experience scores and lower cognitive load compared to No-AI baselines. However, no significant differences were observed in reading speed or comprehension scores". Also the source of the two failure modes named on the slide — over-reliance, and verification that is itself "time-consuming". Small sample, narrow task: cited as a cautionary observation, not a measurement.
### B. Grounded generation: what it fixes and what it does not
- [6] Muhamed, A., Ribeiro, L. F. R., Dreyer, M., Smith, V., & Diab, M. T. (2026). *RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models*. Proceedings of the 19th Conference of the European Chapter of the ACL (EACL 2026), Volume 1: Long Papers, pp. 6811–6856. https://aclanthology.org/2026.eacl-long.321.pdf — accessed 2026-07-27. [type: peer-reviewed] — **The counterweight to the grounded-notebook slides, and the reason the demo asks a question the corpus cannot answer.** Evaluation of over 30 models using 176 perturbation strategies across six categories of informational uncertainty (ambiguity, contradiction, missing information, false premises, granularity mismatch, epistemic mismatch) at three intensity levels. Quoted on the slide: "even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks while exhibiting dangerous over-confidence or over-caution"; some models were "refusing over 60% of answerable queries or confidently answering despite critical information defects"; and "neither scale nor extended reasoning improves performance". Also grounds the framing that refusal is two separable skills — detecting that something is wrong, and correctly categorising *what* is wrong — and that it is "a trainable, alignment-sensitive capability", i.e. it varies by product and by release. **Cohort note (August-2026 revision):** the models evaluated are the 2024–25 generation, and the "neither scale nor extended reasoning improves performance" clause predates the mid-2026 always-on adaptive-thinking releases; the deck now labels it as such and treats the current generation as untested on this benchmark rather than covered by it. The practice the study grounds — test refusal before trusting it — survives either way.
- [7] Imran, S., & Solanky, D. S. (2026). *ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers*. arXiv:2607.11074. https://arxiv.org/abs/2607.11074 — accessed 2026-07-27. [type: peer-reviewed] — Grounds the claim that "chat with this paper" is now a measured task with known failure modes rather than a vibe. A benchmark of 6,211 single-paper question–answer pairs drawn from 494 open-access papers across eight domains (machine learning, public health, education, environmental science, history and humanities, mathematics, psychology, social science) and four question types (lookup, comprehension, multi-hop, adversarial), sourced from OpenAlex's open-access subset. Two findings used on the slides: the benchmark "rewards grounded refusal when the source paper does not support an answer" — establishing refusal as a scored capability, not a bug — and "section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed", i.e. fluency-based judgements cannot separate systems that quote the paper from systems that do not. Preprint; cited for design and direction, not for a leaderboard position.
### C. Structured extraction across many papers
- [8] Gartlehner, G., Kugley, S., Crotty, K., Viswanathan, M., Dobrescu, A., Nussbaumer-Streit, B., et al. (2025). *Artificial Intelligence–Assisted Data Extraction With a Large Language Model: A Study Within Reviews*. Annals of Internal Medicine 178(12), 1763–1771. https://doi.org/10.7326/ANNALS-25-00739 — accessed 2026-07-27. [type: peer-reviewed] — **The best evidence on the deck for what AI-assisted extraction actually costs and buys.** Prospective parallel-group study within reviews (SWAR) across 6 ongoing systematic reviews, blinded adjudicators, using Claude Pro (versions 2.1, 3.0 Opus, 3.5 Sonnet). Concordance between AI-assisted and human-only extraction was 77.2% (95% CI, 76.3% to 78.0%); 9.0% (CI, 8.4% to 9.6%) of AI-assisted extracted items were incorrect against 11.0% (CI, 10.4% to 11.7%) for human-only; major errors 2.5% versus 2.7%; AI assistance saved a median of 41 minutes per study. Authors' conclusion quoted on the slide: "Data extraction assisted by AI may offer a viable, more efficient alternative to human-only methods." Critical framing used on the slide: this is *AI-assisted*, i.e. a human verified every cell — it is not evidence for unsupervised extraction. **Cohort note (August-2026 revision):** the Claude versions used are a 2023–24 cohort; the slide now says so, and the Unreliable-zone placement is argued from the persistence of the accuracy/hallucination tradeoff into the current generation [34], not from these specific percentages.
- [9] Khraisha, Q., Put, S., Kappenberg, J., Warraitch, A., & Hadfield, K. (2024). *Can large language models replace humans in systematic reviews? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages*. Research Synthesis Methods 15(4), 616–626. https://doi.org/10.1002/jrsm.1715 (open preprint: https://arxiv.org/abs/2310.17526) — accessed 2026-07-27. [type: peer-reviewed] — The "human-out-of-the-loop" contrast to [8]. Pre-registered evaluation of GPT-4 on title/abstract screening, full-text review and data extraction across literature types and languages. Quoted on the slide: "Although GPT-4 had accuracy on par with human performance in most tasks, results were skewed by chance agreement and dataset imbalance. After adjusting for these, there was a moderate level of performance for data extraction, and — barring studies that used highly reliable prompts — screening performance levelled at none to moderate." Conclusion used on the slide: "substantial caution should be used if LLMs are being used to conduct systematic reviews", but "for certain systematic review tasks delivered under reliable prompts, LLMs can rival human performance". Together [8] and [9] make the deck's point precisely: the human in the loop is the variable that moves the number.
- [11] Walters, W. H., & Wilder, E. I. (2023). *Fabrication and errors in the bibliographic citations generated by ChatGPT*. Scientific Reports 13, 14045. https://www.nature.com/articles/s41598-023-41032-5 — accessed 2026-07-27 (carried forward from Sessions 1 and 2). [type: peer-reviewed] — Carried forward so the series quotes one number for one claim. Across 636 citations in 84 generated literature reviews, "55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated", and "43% of the real (non-fabricated) GPT-3.5 citations but just 24% of the real GPT-4 citations include substantive citation errors". Session 3 uses only the second half: a citation can be real and still misdescribe what the source says — which is exactly the residual failure that source-grounding does not remove.
- [12] Lieberum, J.-L., Toews, M., Metzendorf, M.-I., Heilmeyer, F., Siemens, W., Haverkamp, C., Böhringer, D., Meerpohl, J. J., & Eisele-Metzger, A. (2025). *Large language models for conducting systematic reviews: on the rise, but not yet ready for use — a scoping review*. Journal of Clinical Epidemiology 181, 111746. https://www.jclinepi.com/article/S0895-4356(25)00079-4/fulltext — accessed 2026-07-27 (carried forward from Sessions 1 and 2). [type: peer-reviewed] — Grounds the placement of "extract structured data from many papers" in the **Unreliable** zone of Session 1's Capability/Failure Map: 37 included articles covering 10 of 13 defined systematic-review steps, with data extraction among the most-studied (30% of studies); "about half of the studies found LLMs promising, a quarter were neutral, and one-fifth found them not promising"; conclusion that "LLMs should only be used with caution and under human supervision".
- [13] Flemyng, E., Noel-Storr, A., Macura, B., Gartlehner, G., Thomas, J., Meerpohl, J. J., Jordan, Z., Minx, J., Eisele-Metzger, A., Hamel, C., Jemioło, P., Porritt, K., & Grainger, M. (2025). *Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence 2025*. Cochrane Database of Systematic Reviews, Editorial ED000178. https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.ED000178/full — accessed 2026-07-27 (carried forward from Session 2). [type: policy] — The governing rule for anyone taking an AI-built extraction table into a formal synthesis, quoted on the extraction-accuracy slide and in the handout checklist: authors are "ultimately responsible" for the synthesis including the decision to use AI; AI may be used "as long as you can demonstrate that it will not compromise the methodological rigour or integrity" of the synthesis; "AI and automation in evidence synthesis should be used with human oversight"; and "Any use of AI or automation that makes or suggests judgements should be fully and transparently reported in the evidence synthesis report."
- [34] Anthropic (2026). *System Card: Claude Opus 5* (24 July 2026). https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf — accessed 2026-08-21. [type: primary-doc] — Added in the August-2026 revision to carry the extraction slide's zone re-argument. A vendor self-evaluation, used only for facts about the vendor's own model and labelled as such on the slide. The load-bearing quotes: "We found a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure. The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall"; and on the public split of AA-Omniscience, "the Claude Opus 5's accuracy is 11% higher than Opus 4.8, but its rate of hallucinations is also 6% higher." This is why the deck keeps "extract structured data" in the Unreliable zone without leaning on 2024's percentages: the guess-more-versus-abstain-more tradeoff is structural and the vendor documents it moving in the wrong direction even as accuracy improves.
### D. Primary documentation — grounded notebooks (Gemini Notebook, formerly NotebookLM)
- [14] Google (2026). *NotebookLM is now Gemini Notebook* (announcement, 16 July 2026, Josh Woodward, VP Google Labs). https://blog.google/innovation-and-ai/products/gemini-notebook/notebooklm-gemini-notebook/ — accessed 2026-07-27. [type: primary-doc] — The freshness fact the deck opens the grounded-notebook section with, cited to the vendor's own announcement of its own product name: "Today, we're renaming NotebookLM to Gemini Notebook. It remains a standalone product focused on being your premier research tool, but it will now do more across the Google ecosystem, including inside the Gemini app and Google Search." Also the source of the adoption figure quoted on the slide as vendor self-reported — "more than 30 million people and over 600,000 organizations are using it" — and of the July 2026 capability change: every notebook is getting "a secure cloud computer" that "allows Gemini Notebook to write and execute code natively, helping you conduct complex data analysis grounded in your sources", available first to Google AI Ultra and Workspace AI Ultra/Expanded Access customers and rolling out to Pro users on the web. **Verification note (2026-08-21):** the rename, the standalone-product framing and the continued validity of existing links were re-confirmed on Google's Workspace Updates post of 16 July 2026 (https://workspaceupdates.googleblog.com/2026/07/notebooklm-now-gemini-notebook.html, accessed 2026-08-21); the secure-cloud-computer code-execution capability for AI-Pro subscribers was corroborated the same day by 9to5google's report of the launch, alongside the "over 30 million people and 600,000+ organizations" usage figure.
- [15] Google (2026). *Learn about Gemini Notebook — Computer*. Gemini Notebook Help. https://support.google.com/gemininotebook/answer/16164461 — accessed 2026-07-27. [type: primary-doc] — Primary definition used on the "what a grounded notebook is" slide: "Gemini Notebook is an AI-powered research assistant designed to help you refine and organize your ideas"; you can "Easily upload PDFs, websites, YouTube videos, audio files, Google Docs, or Google Slides"; "Chat with your notebook to get grounded information based on your sources with clear in-line citations for accuracy, transparency, and trust"; and transform sources "into approachable formats such as study guides, briefings, audio overviews, mind maps, and more". Also grounds the three documented reasons a notebook fails to answer — safety flags, unclear phrasing, and "Information not in sources" — and the retrieval mechanic stated in the vendor's own words: "When your notebook contains many sources, Gemini Notebook retrieves the most relevant information based on your question first, then builds a response with this information."
- [16] Google (2026). *Use chat in Gemini Notebook*. Gemini Notebook Help. https://support.google.com/gemininotebook/answer/16179559 — accessed 2026-07-27. [type: primary-doc] — Grounds the citation mechanics demonstrated live in the demo: "Gemini Notebook uses direct quotes, text, and images straight from your sources as citations to answer your questions and perform actions. These citations help you check the accuracy of the response. You can hover over any citation to get the full quoted text right away. If you select a citation, Gemini Notebook automatically navigates to the location of the quote, so you can easily view it in context." Also grounds the corpus-scoping move in Demo step 3: "you can use the checkbox on each source to include or exclude certain sources the model should use to answer your question"; and the refusal behaviour the demo elicits — chat responses "only use data from your sources", and out-of-scope requests return "Gemini Notebook can't answer this question".
- [17] Google (2026). *Add or discover new sources for your notebook — Computer*. Gemini Notebook Help. https://support.google.com/gemininotebook/answer/16215270 — accessed 2026-07-27. [type: primary-doc] — Grounds the corpus-building slide and two of the handout's setup steps. Definitional: "A source is a copy or auto-synced version of the source document you import or upload to the app. When you use Gemini Notebook, the model uses the sources you upload to answer your questions or complete your requests." Documented limitations quoted in the handout: "Gemini Notebook does not import footnotes or comments from Google files" and "Importing audio files from Drive is not supported". Also the citation-granularity caveat used on the critical-reading slide: "If the source content is too short, Gemini Notebook references the entire document without citing individual text from your source." Finally, it documents the Deep Research import path — Deep Research "can automatically browse up to hundreds of websites on your behalf" and the results can be selected and imported as sources — which is the bridge from Session 2's Tier 3 agents into a Session 3 notebook.
- [18] Google (2026). *Frequently asked questions*. Gemini Notebook Help. https://support.google.com/gemininotebook/answer/16269187 — accessed 2026-08-21 (re-fetched; all figures below verified unchanged from the 2026-07-27 capture). [type: primary-doc] — The hard numbers on the corpus-limits slide and in the handout's setup checklist, all first-party: "Gemini Notebook answers questions based on the information provided in your uploaded sources. If the answer isn't in the source material, it won't provide a response"; "Get 100 notebooks, with up to 50 sources each and 500,000 words each"; "daily limits of 50 chat queries and 3 audio generations"; "The current limit is 500,000 words per source or up to 200MB for local uploads. There's no page limit"; import fails if a source exceeds those limits or if "Your original PDF file is copy-protected"; and the prompt-composition rule that matters for reproducibility — "Sources: Always used in either the entire set or the subset you select", "Notes: Only when you specifically select it", "Conversation history: Used to generate responses". Higher limits are stated to require an upgraded plan.
### E. Primary documentation — extraction, PDF chat and reference managers
- [19] Elicit (2026). *Improve column results*. Elicit Help Center. https://support.elicit.com/en/articles/14758163-improve-column-results — accessed 2026-07-27. [type: primary-doc] — **The single most useful primary-doc fact in the session**, and the vendor's own admission that grounds the "abstract trap" slide alongside [4]: "When Elicit has access to the PDF of the paper, Elicit can search for information from the entire paper. If Elicit does not have access to the PDF for the paper, Elicit can only pull information from the abstract. If the information you're looking for is not in the abstract, we may not be able to extract it for you. Elicit typically has access to open-access papers." Also grounds the three documented remedies used in the handout — the "Has PDF" filter toggle, the browser extension for institutional full text, and uploading your own PDFs or importing them from Zotero collections.
- [20] Elicit (2026). *Create and save columns in Elicit*. Elicit Help Center. https://support.elicit.com/en/articles/14758162-create-and-save-columns-in-elicit — accessed 2026-07-27. [type: primary-doc] — Grounds the structure of the session's Extraction Column Template artefact, taken from a working tool's own guidance rather than invented: a column has a **Name** ("a single phrase describing what information you want to extract from the papers in your table, e.g. 'Main or primary outcome measured in the study'") and **Instructions** that "tell Elicit how to get the information you want", including the desired format and an explicit level of precision ("Be as precise as possible (i.e. '50 hours' is better than '2 days'.)", "'mean ± standard deviation'"). Also grounds the two constrained answer types the handout recommends for screening — Yes/No/Maybe columns, and "Specified" multiple-choice columns where "Elicit will choose answers from your list of specified answers".
- [21] Elicit (2026). *Systematic Reviews in Elicit*. Elicit Help Center. https://support.elicit.com/en/articles/14759154-systematic-reviews-in-elicit — accessed 2026-07-27. [type: primary-doc] — Grounds the verification affordance that the deck argues is non-negotiable in an extraction tool: "Click on any cell in the finished data extraction table to view supporting quotes from the paper so that you can easily check the AI-generated answers for accuracy." Also documents the pilot-then-run workflow the handout recommends ("You'll first define extraction columns on a subset of papers to make sure your columns are functioning as you desire"), and the tier-dependence of figure/table reading ("For Scale and Enterprise users, extractions can pull from figures, tables, charts, and diagrams in addition to text"). The vendor's efficiency claim on the same page — "save up to 80% of the time on a systematic review without sacrificing accuracy" — is **not** used anywhere in the deck or handout; it is a vendor performance claim with no independent evaluation attached.
- [22] SciSpace (2026). *Chat with any PDF | AI-powered Research Paper & PDF Summarizer*. https://scispace.com/chat-pdf — accessed 2026-07-27. [type: primary-doc] — Primary description for the SciSpace row of the reading-tool table, in the vendor's own words: "Upload any PDF to SciSpace Chat PDF, ask a question, and get concise, citation-linked answers, summaries, and follow-ups in seconds — free tier, 256-bit encrypted, no data training, supports 75+ languages"; "Get answers backed by citations from specific sections of the PDF"; section-wise paper summaries; explanations of maths, equations, tables and figures; and chat across multiple PDFs.
- [23] SciSpace (2026). *AI-Powered Data Extraction from Research PDFs*. https://scispace.com/extract-data — accessed 2026-07-27. [type: primary-doc] — Grounds the SciSpace extraction cell and the comparison-table mechanics: "SciSpace Data Extractor identifies tables, stats and citations in research PDFs, summarises key findings and exports clean data to CSV, Excel or RIS"; "Apply 50 parameters to PDFs or research papers to compare at once"; and the documented ceiling stated in the vendor's own FAQ, "Currently you can add 50 columns including the default suggestions and those you create manually." The on-page competitor tick-box comparison against Elicit and Consensus is a vendor marketing artefact and is **not** used as evidence anywhere in this session.
- [24] Paperpal (2026). *AI Academic Writing Tool — Comprehensive AI Research Assistant* (page dated 14 July 2026; Cactus Communications). https://paperpal.com/ — accessed 2026-07-27. [type: primary-doc] — Included so the reading-tool table covers the publisher-services corner of the market as well as the start-up corner. Vendor's own description of the relevant feature: "AI Chat with PDFs — Stop manually cross-referencing papers. Upload your sources, ask questions, and get key findings with citations. Compare across documents to build your literature review in a fraction of the time." Also documents an adjacent feature the deck mentions once in passing, Reference Check, which the vendor describes as fixing "citation issues, broken links, and AI-hallucinated references". Marketing claims on the same page (journal counts, awards) are not used.
- [25] SCiNiTO (2026). *PDF Analysis*. SCiNiTO user guide. https://www.scinito.ai/help/pdf-analysis — accessed 2026-07-27. [type: primary-doc] — SCiNiTO's row in the reading-tool table, described from its own help centre and listed as one option among peers with no endorsement: "Upload any scholarly article and get structured, AI-powered insights in minutes — not hours of reading"; the documented workflow is "Upload a PDF, browse the structured insights, and ask follow-up questions about the paper"; and an Analysis history that lets you "Find every PDF you've previously analyzed and return to its insights any time." Session 2's entry for SCiNiTO's search and chat behaviour is not repeated here.
- [26] Zotero (2026). *Plugins for Zotero*. Zotero Documentation. https://www.zotero.org/support/plugins — accessed 2026-07-27. [type: primary-doc] — Grounds the reference-manager slide's central, deliberately unglamorous point, in Zotero's own words: "We don't currently provide a list of available plugins, but most plugins are announced and discussed in the Zotero Forums. An official plugin directory is planned." And the security warning quoted verbatim on the slide and in the handout checklist: "Be aware that plugins have full access to your Zotero and your computer. You should only install plugins from developers you trust." Also documents what *is* bundled — word-processor plugins for Word, LibreOffice and Google Docs, and browser connectors — which is the honest answer to "does Zotero have AI built in": the documentation describes no first-party AI feature, so any AI in a Zotero workflow arrives as third-party code under that warning.
- [27] Zotero (2026). *Zotero PDF Reader and Note Editor*. Zotero Documentation. https://www.zotero.org/support/pdf_reader — accessed 2026-07-27. [type: primary-doc] — Grounds the "the non-AI half of the workflow still matters" slide and the handout's integration steps: annotations added to notes "will automatically include links back to the PDF page as well as citations that you can later add to a Word, LibreOffice, or Google Docs document with one of the word processor plugins"; you can "create a child note from all annotations in a PDF" or "a standalone note with annotations from multiple items"; and clicking an annotation and choosing "Show on Page" reopens the original PDF at the page where the annotation was made. This is the manual, fully auditable version of the quote-to-source link that a grounded notebook automates — the slide uses it to show that the citation-anchoring idea predates the AI feature.
### F. The demo corpus — five open-access papers on one question
*These five are the Session 3 demo corpus (see `demo-script.md`). They are also cited on the slide that previews the extraction table, so their numbers appear here. All five were verified against Crossref on 2026-07-27, and open-access status was verified against Unpaywall on the same date. All five DOIs were re-resolved against the Crossref API on 2026-09-02 (registered, titles matching); the Unpaywall open-access check was not repeated.*
- [28] Davis, P. M., Lewenstein, B. V., Simon, D. H., Booth, J. G., & Connolly, M. J. L. (2008). *Open access publishing, article downloads, and citations: randomised controlled trial*. BMJ 337, a568. https://doi.org/10.1136/bmj.a568 — accessed 2026-07-27. [type: peer-reviewed] — Demo corpus paper 1, and the corpus's only randomised design. 11 journals published by the American Physiological Society, 1,619 research articles and reviews randomly assigned at online publication to open access or subscription access. Results quoted in the extraction table: "89% more full text downloads (95% confidence interval 76% to 103%), 42% more PDF downloads (32% to 52%), and 23% more unique visitors (16% to 30%), but 24% fewer abstract downloads (−29% to −19%)" in the first six months; "Open access articles were no more likely to be cited than subscription access articles in the first year after publication" — 59% (146/247) of open-access articles cited at 9–12 months against 63% (859/1372) of subscription articles. Conclusion: "The citation advantage from open access reported widely in the literature may be an artefact of other causes." Open access (hybrid, CC-BY) per Unpaywall.
- [29] Ottaviani, J. (2016). *The Post-Embargo Open Access Citation Advantage: It Exists (Probably), It's Modest (Usually), and the Rich Get Richer (of Course)*. PLOS ONE 11(8), e0159614. https://doi.org/10.1371/journal.pone.0159614 — accessed 2026-07-27. [type: peer-reviewed] — Demo corpus paper 2, chosen for its middle-of-the-road effect size and for a teaching detail: it carries a published Correction (PLOS ONE 11(10), e0165166, 21 October 2016) that a notebook built only from the article PDF will not know about. Abstract figure used in the extraction table: "this study addresses those factors and shows that an open access citation advantage as high as 19% exists, even when articles are embargoed during some or all of their prime citation years", with above-median articles gaining more. Design and sample, read from the Methodology section and used in the demo-corpus table on slide 23: a random selection of 3,850 peer-reviewed and review articles from Deep Blue, the University of Michigan institutional repository (original publication dates 1990–2013), matched against the 89,895 corresponding articles that remained closed in the same journal issues; crucially "None of the OA articles were self-selected; authors did not choose to deposit the articles in question in Deep Blue, since they were opened via blanket licensing agreements between the publishers and the library" — which is what makes this an observational study with a genuine matched control rather than a self-selection artefact. Gold open access, CC-BY.
- [30] Piwowar, H., Priem, J., Larivière, V., Alperin, J. P., Matthias, L., Norlander, B., Farley, A., West, J., & Haustein, S. (2018). *The state of OA: a large-scale analysis of the prevalence and impact of Open Access articles*. PeerJ 6, e4375. https://doi.org/10.7717/peerj.4375 — accessed 2026-07-27. [type: peer-reviewed] — Demo corpus paper 3, the large-scale observational study. Three samples of 100,000 articles each, drawn from Crossref DOIs, recent Web of Science records, and articles viewed by Unpaywall users; oaDOI determined open-access status for 67 million articles. Figures used in the extraction table: "at least 28% of the scholarly literature is OA (19M in total)"; 45% for the most recent year analysed (2015); and the headline the demo's first question is designed to surface — "accounting for age and discipline, OA articles receive 18% more citations than average, an effect driven primarily by Green and Hybrid OA". Gold open access, CC-BY.
- [31] Langham-Putrow, A., Bakker, C., & Riegelman, A. (2021). *Is the open access citation advantage real? A systematic review of the citation of open access and subscription-based articles*. PLOS ONE 16(6), e0253129. https://doi.org/10.1371/journal.pone.0253129 — accessed 2026-07-27. [type: peer-reviewed] — Demo corpus paper 4, the systematic review that makes the corpus disagree with itself in a structured way. Systematic search of 17 databases, protocol registered in OSF; 5,744 items retrieved, 134 included. Results used in the extraction table and in the demo's "does it surface the disagreement?" test: "64 studies (47.8%) confirmed the existence of OACA, while 37 (27.6%) found that it did not exist, 32 (23.9%) found OACA only in subsets of their sample, and 1 study (0.8%) was inconclusive"; and the detail most likely to be flattened by a summary — "In the critical appraisal of the included studies, 3 were found to have an overall low risk of bias. Of these, one found that an OACA existed, one found that it did not, and one found that an OACA occurred in subsets." Gold open access, CC-BY.
- [32] Huang, C.-K., Neylon, C., Montgomery, L., Hosking, R., Diprose, J. P., Handcock, R. N., & Wilson, K. (2024). *Open access research outputs receive more diverse citations*. Scientometrics 129(2), 825–845. https://doi.org/10.1007/s11192-023-04894-0 — accessed 2026-07-27. [type: peer-reviewed] — Demo corpus paper 5, chosen because it changes the *outcome variable* rather than the answer, which is what makes the comparison table worth building. Large-scale bibliographic analysis 2010–2019 finding "a robust association between open access and increased diversity of citation sources by institutions, countries, subregions, regions, and fields of research, across outputs with both high and medium–low citation counts", with repository-based open access showing "a stronger effect than open access via publisher platforms". It also cites [28]'s randomised evidence against itself — "A set of narrowly defined randomised control trials finds no effect" — which is the honest scholarly move the demo asks the notebook to reproduce. Open access (hybrid, CC-BY).
<!--
Counts for this session (checked 2026-07-27; revised 2026-08-21):
total entries = 34 (target >= 15)
peer-reviewed/preprint = 18 ([1] [2] [3] [4] [5] [6] [7] [8] [9] [10] [11] [12] [28] [29] [30] [31] [32] [33]; target >= 5)
policy = 1 ([13])
primary-doc = 15 ([14]–[27] [34])
news = 0
Sources [11], [12] and [13] are deliberately shared with Sessions 1 and 2 so the
series quotes the same numbers for the same claims across sessions.
Sources [28]–[32] are the demo corpus and are also this session's canonical
artefact for later sessions (Session 6 may reuse the corpus).
-->
### Note on vendor documentation and the source standard
Fifteen entries ([14]–[27], [34]) are vendor documentation. Each is used **only** for facts about that vendor's own product — what it does, what it costs, what it refuses to do, what its limits are — which the series' source standard explicitly permits. No comparative or performance claim about *summarisation or extraction accuracy* in this deck rests on vendor material:
- Every claim about how accurate summarisation or extraction is rests on [1]–[12].
- Two vendor performance claims were found during research and **deliberately excluded**: Elicit's "save up to 80% of the time on a systematic review without sacrificing accuracy" [21], and SciSpace's benchmark claim that its Deep Review returned 26.3 highly relevant papers per query against Elicit's 13.0. Neither has an independent evaluation attached; neither appears on any slide or in the handout.
- One vendor figure *is* quoted, on the grounded-notebook slide: "more than 30 million people and over 600,000 organizations" using Gemini Notebook [14]. It is labelled on-slide as vendor self-reported adoption data, per Ehsan's 2026-07-27 sign-off on first-party figures, and nothing in the argument depends on it.
- The August-2026 revision adds a second labelled vendor figure, on the extraction-accuracy slide: Anthropic's own Opus 5 system-card statement that accuracy rose 11% while hallucinations rose 6% [34]. It is attributed to "Anthropic's own Opus 5 card" in the slide text, is a self-report about the vendor's own model (the direction least flattering to the vendor), and is used to argue that the guess-vs-abstain tradeoff persists — not as an accuracy benchmark.
### Dropped during research (and why)
- **"NotebookLM has a response-level hallucination rate of around 13%, compared with around 40% for general LLMs."** This figure circulates widely and appears in several AI-summary sites and consultancy write-ups. It traces to no primary study that could be located, is attributed vaguely to "testing in journalistic workflows", and no methodology is available. **Dropped entirely.** The grounded-notebook slides make the qualitative claim (grounding removes fabrication of sources, not misreading of them) from the vendor's own documented behaviour [15][16][18] and the peer-reviewed refusal evidence [6][7].
- **"Hedging language decreased by 22.8% (1.88 → 1.45 occurrences per 1,000 words), p < 10⁻⁶³."** A precise-looking figure returned by search without a resolvable primary source; it does not appear in [1], which measures generic-vs-quantified conclusions rather than hedge frequency. Dropped; the hedging point is made from [4], whose finding was read at the source.
- **Zotero AI-plugin roundups and rankings** (`papersflow.ai`, `citationstyler.com`, and similar "7 best Zotero AI plugins in 2026" listicles). Useful for discovering that the plugin ecosystem exists; useless as evidence, and several are affiliate-shaped. The reference-manager slide therefore names no third-party plugin at all and instead quotes Zotero's own security warning [26], which is the durable advice.
- **NotebookLM/Gemini Notebook source-limit figures by paid plan** ("50 / 100 / 300 / 600 sources for Standard / Plus / Pro / Ultra"). Widely repeated across secondary sites; only the free-tier numbers (100 notebooks, 50 sources, 500,000 words, 50 chat queries/day, 3 audio generations/day) could be confirmed in Google's own help centre [18], so only those appear on the slide, explicitly labelled as the documented free-tier limits.
- **Davis, P. M. (2011), *Open access, readership, citations: a randomized controlled trial of scientific journal publishing*, FASEB J 25(7), 2129–2134** (doi:10.1096/fj.11-183988). Originally selected as demo corpus paper 2 — it is the larger, longer-follow-up sibling of [28]. Unpaywall reports it as **closed** (checked 2026-07-27), which breaks the demo's requirement that every attendee can rebuild the corpus themselves. Replaced with [29]. The paper is still worth naming aloud in the demo as the study [32] refers to when it says randomised trials find no effect.
- **General "RAG reduces hallucination by X%" benchmark claims** from vendor leaderboards and evaluation-platform marketing pages. Not transferable to a specific product's behaviour on a specific corpus, and not peer-reviewed. The deck instead uses [6] and [7], which measure the two things a researcher actually cares about: does it cite the passage it claims to, and does it refuse when it should.
- **The retracted comparative evaluation of Elicit, SciSpace and Consensus** flagged in Session 2's `sources.md`. It resurfaced in this session's searches on "structured data extraction from papers". Dropped again, for the same reason.
---
*AI for Researchers · Session 3: Reading, Notes & Knowledge Management · Landscape as of August 2026*