AI for Researchers

Session 3: Reading, Notes & Knowledge Management

Grounded notebooks, structured extraction, and the nuance a summary quietly drops

{{Presenter Name}}

Landscape as of August 2026

2-Minute Recap

The Series So Far


Today (Session 3): the shortlist becomes a corpus — and we find out what a summary of it leaves out.

Learning Objectives

By the End of This Session You Will Be Able To…


  1. Explain how AI-assisted paper Q&A / PDF-chat tools work and where summarization flattens methodological nuance.
  2. Build a source-grounded notebook (Gemini Notebook-style, formerly NotebookLM) from a small set of papers.
  3. Run structured queries against a grounded notebook to extract methods, effect sizes, and populations.

Objectives 4–6 continue on the next slide.

Learning Objectives

…And You Will Also Be Able To


  1. Apply AI-assisted extraction to build a cross-paper comparison table.
  2. Evaluate when an AI summary requires returning to the primary source rather than trusting the synthesis.
  3. Integrate AI reading tools into a reference-manager workflow (e.g., Zotero).

Section 01

What "Chat With This Paper" Actually Does

Before we trust an answer about a PDF, it is worth knowing which part of the PDF the tool was allowed to look at.

Literacy Foundation

The Mechanism, in Four Steps


  1. Index. Your PDF is broken into passages the tool can look up individually. [16]
  2. Retrieve. Your question pulls back the passages that look most relevant. [15]
  3. Generate. The model writes an answer using those passages as its evidence.
  4. Cite. Good tools attach the quoted passage, so you can click back to it. [16]

Google states the retrieval step plainly: the notebook "retrieves the most relevant information based on your question first, then builds a response with this information" [15]. Step 2 is where most failures are born, and it is the step you cannot see. One 2026 update: million-token flagships (Session 1) can now ingest a small corpus whole — no retrieval step — but that removes lost passages, not misreading, and at notebook scale retrieval is still what happens.

Tier 2 — Research-Specific Tools

Reading and Extraction Tools, From Their Own Documentation


Tool What it is built to do How it shows its evidence Documented limit
Gemini Notebookformerly NotebookLM [14] Source-bound Q&A over a corpus you upload [15] In-line citations; hover for the quote, click to jump to it [16] Free tier: 50 sources per notebook, 50 chat queries/day [18]
Elicit Extraction tables: one column per data point across many papers [20] Click any cell to see the supporting quote from the paper [21] Without the PDF it "can only pull information from the abstract" [19]
SciSpace PDF chat plus a data extractor over uploaded papers [22] "Answers backed by citations from specific sections of the PDF" [22] Maximum 50 extraction columns per table [23]
Paperpal Chat across uploaded sources, inside a writing workflow [24] "Get key findings with citations" [24] Writing-workflow scope; corpus limits not documented [24]
SCiNiTO Structured insights from an uploaded article, plus follow-ups [25] Analysis history keeps every prior PDF and its insights [25] Per-article analysis; cross-corpus use not documented [25]
Zotero + plugins Your library is the corpus; AI arrives as third-party code [26] Annotations carry page links and citations into notes [27] "Plugins have full access to your Zotero and your computer" [26]
Landscape as of August 2026. Every cell is taken from that tool's own current documentation; no cell is a comparative judgement. Row order is not a ranking — no endorsement implied.

Critical Literacy

The Abstract Trap


Section 02

Where Summaries Flatten Nuance

Four independent lines of evidence, all pointing at the same thing: the qualifier is the first casualty.

The Evidence — 1 of 4

Summaries Systematically Overgeneralise


26–73% of cases overgeneralised, for three of the ten models tested [1]
OR 1.90 Prompting for accuracy made overgeneralisation more likely [1]
OR 4.85 vs. expert-written summaries of the same research [1]

Numbers above: the 2025 ten-model study [1]. No study yet tests the August-2026 cohort on overgeneralisation — an open gap, not an all-clear.

The Evidence — 2 of 4

Ask the Authors What Went Missing


Qualitative study, 20 interviews, reported as a refereed conference poster; summaries were ChatGPT-4 (2025). Cited for which content disappears, never as a rate.

The Evidence — 3 of 4

Reading the Summary Instead Costs You Comprehension


d = 0.83 How much the summary tool worsened comprehension for strong readers [3]
d = 0.45 How much it helped weaker readers — the smallest gain of four tools [3]

Standardised test passages, not journal articles. Take the direction; do not quote the effect size for research reading.

The Evidence — 4 of 4

Even Experts Need a Checklist


Series Artefact

The Summary-Trust Triage: When to Reopen the Paper


  1. Is there a number in it? Effect sizes, n, CIs — go to the source. Always.
  2. Has a qualifier vanished? No population, no design, no boundary condition = a generic. [1]
  3. Is it a limitation or a method? The two things summaries drop first. [2]
  4. Does a claim carry no citation? Unsupported prose in a grounded tool is a red flag, not a stylistic choice. [16]
  5. Will you cite it? Then the summary was never the evidence. The paper is. [11]

Rule of thumb: use summaries to decide what to read and what to compare — never as the thing you cite.

Section 03

Grounded Notebooks

A corpus you chose, a model that may only answer from it, and a citation on every sentence. What that buys — and what it does not.

Grounded Notebooks

"Grounded" Is a Constraint, Not a Compliment


Vendor self-reported adoption, quoted as such: Google states "more than 30 million people and over 600,000 organizations" use it [14]. Nothing on this slide depends on that number.

Critical Literacy

Refusal Is a Capability — and It Is Unreliable


< 50% Refusal accuracy on multi-document tasks, even for frontier models [6]
> 60% Answerable questions refused, by some models at the other extreme [6]

Grounded Notebooks

The Corpus Is the Method


Section 04

From Reading to a Table

Twenty papers, one grid: methods, populations, effect sizes. The part of AI-assisted reading with the best evidence behind it — and the strictest conditions.

Structured Extraction

Stop Asking for a Summary. Ask for a Column.


This is the same discipline as a systematic-review extraction form. The tool changes; the form does not.

The Evidence

How Good Is It? It Depends Entirely on the Human


9.0% Items incorrect, AI-assisted extraction [8]
11.0% Items incorrect, human-only extraction [8]
41 min Median time saved per study [8]

Series Artefact

The Extraction Column Template


Name: one phrase — "Primary outcome measured in the study".
Instruction: where to look, what counts, what format.
Precision: "Be as precise as possible — '50 hours' is better than '2 days'." [20]
Abstain rule: "If the paper does not state this, write NOT REPORTED. Do not infer it."

Today's Demo Corpus

Five Open-Access Papers, One Question: Does Open Access Raise Citations?


Paper Design Population / sample Headline result
Davis et al. 2008BMJ 337, a568 Randomised controlled trial 1,619 articles, 11 physiology journals +89% full-text downloads; no citation advantage at one year [28]
Ottaviani 2016PLOS ONE 11(8) Observational, matched control 3,850 opened vs 89,895 closed articles Advantage "as high as 19%"; bigger for above-median articles [29]
Piwowar et al. 2018PeerJ 6, e4375 Large-scale observational 3 samples × 100,000 articles OA articles receive 18% more citations than average [30]
Langham-Putrow et al. 2021PLOS ONE 16(6) Systematic review 134 studies from 5,744 screened 47.8% found an advantage, 27.6% found none, 23.9% in subsets [31]
Huang et al. 2024Scientometrics 129 Large-scale observational Global outputs, 2010–2019 Different outcome: OA is cited from more places [32]
Landscape as of August 2026. All five are open access with resolvable DOIs — re-verified against Crossref 2026-09-02; open-access status verified against Unpaywall 2026-07-27. The full list is in this session's demo script so you can rebuild the corpus yourself.

Critical Literacy

The Column That Decides the Answer Is "Design"


Section 05

Notes That Survive the Tool

Notebooks come and go. Your reference manager is where the reading has to land if you want it in two years' time.

Knowledge Management

Where the AI Fits in a Zotero Workflow


What Zotero already does

  • Highlights become notes that carry a link back to the exact page. [27]
  • "Show on Page" reopens the PDF where you made the annotation. [27]
  • Those citations flow into Word, LibreOffice or Google Docs. [27]
  • This is quote-to-source grounding — done by hand, and fully auditable.

What the plugins add — and cost

  • AI in Zotero arrives as third-party code, not a built-in feature. [26]
  • Zotero's own warning: "plugins have full access to your Zotero and your computer". [26]
  • There is still no official plugin directory — "an official plugin directory is planned". [26]
  • Lower-risk path: export the PDFs out to the reading tool instead. [19]

Systematic Prompting

CRIT, Applied to Reading


One paper, in PDF chat

Name the section, ask one thing, and demand the quote.

  • C: which paper, which section.
  • I: quote the sentence you used; say NOT REPORTED if absent.
  • T: one field at a time — then re-ask.

Many papers, in a notebook

Ask for a grid, per source, with disagreements kept intact.

  • C+R: the corpus, the field, the register.
  • I: answer per source; never merge sources into one claim. [5]
  • T: then ask what the corpus cannot answer. [6]

Live Demo & Takeaways

What We Do Next — and What You Take Home


References

References (1–16)


  1. Peters & Chin-Yee (2025). Generalization bias in large language model summarization of scientific research. R. Soc. Open Sci. 12(4), 241776. doi.org/10.1098/rsos.241776 — accessed 2026-08-21
  2. Hammond & Newell (2025). Reading between the GenAI lines. ASCILITE Publications (refereed poster). doi.org/10.65106/apubs.2025.2719 — accessed 2026-07-27
  3. Etkin, Etkin, Carter & Rolle (2025). Differential effects of GPT-based tools on comprehension of standardized passages. Frontiers in Education 10, 1506752. doi.org/10.3389/feduc.2025.1506752 — accessed 2026-07-27
  4. Thelwall (2026). Will AI be overconfident about academic research findings when reliant on abstracts? arXiv:2605.27392. arxiv.org/abs/2605.27392 — accessed 2026-07-27
  5. Martin, Humphreys, Brown, Leckey & Kaur (2026). An Expert Schema for Evaluating LLM Errors in Scholarly QA Systems. arXiv:2602.21059. arxiv.org/abs/2602.21059 — accessed 2026-07-27
  6. Muhamed, Ribeiro, Dreyer, Smith & Diab (2026). RefusalBench: Generative Evaluation of Selective Refusal in Grounded LMs. EACL 2026. aclanthology.org/2026.eacl-long.321.pdf — accessed 2026-07-27
  7. Imran & Solanky (2026). ResearchQA: Benchmarking Citation-Grounded QA on Scientific Papers. arXiv:2607.11074. arxiv.org/abs/2607.11074 — accessed 2026-07-27
  8. Gartlehner et al. (2025). AI-Assisted Data Extraction With a Large Language Model: A Study Within Reviews. Ann Intern Med 178(12). doi.org/10.7326/ANNALS-25-00739 — accessed 2026-07-27
  9. Khraisha, Put, Kappenberg, Warraitch & Hadfield (2024). Can LLMs replace humans in systematic reviews? Res. Synth. Methods 15(4). doi.org/10.1002/jrsm.1715 — accessed 2026-07-27
  10. Xu et al. (2026). Comparing structural and conversational AI summarisation for academic reading. Interact. Learn. Environ. doi.org/10.1080/10494820.2025.2604649 — accessed 2026-07-27
  11. Walters & Wilder (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci. Rep. 13, 14045. nature.com/articles/s41598-023-41032-5 — accessed 2026-07-27
  12. Lieberum et al. (2025). LLMs for conducting systematic reviews: not yet ready for use. J. Clin. Epidemiol. 181, 111746. jclinepi.com/article/S0895-4356(25)00079-4 — accessed 2026-07-27
  13. Flemyng et al. (2025). Position statement on AI use in evidence synthesis (Cochrane, Campbell, JBI, CEE). Cochrane DSR, ED000178. cochranelibrary.com/cdsr/doi/10.1002/14651858.ED000178 — accessed 2026-07-27
  14. Google (2026). NotebookLM is now Gemini Notebook (announcement, 16 July 2026). blog.google/innovation-and-ai/products/gemini-notebook/notebooklm-gemini-notebook/ — accessed 2026-07-27
  15. Google (2026). Learn about Gemini Notebook. Gemini Notebook Help. support.google.com/gemininotebook/answer/16164461 — accessed 2026-07-27
  16. Google (2026). Use chat in Gemini Notebook. Gemini Notebook Help. support.google.com/gemininotebook/answer/16179559 — accessed 2026-07-27

References

References (17–34)


  1. Google (2026). Add or discover new sources for your notebook. Gemini Notebook Help. support.google.com/gemininotebook/answer/16215270 — accessed 2026-07-27
  2. Google (2026). Frequently asked questions. Gemini Notebook Help. support.google.com/gemininotebook/answer/16269187 — accessed 2026-08-21
  3. Elicit (2026). Improve column results. Elicit Help Center. support.elicit.com/en/articles/14758163 — accessed 2026-07-27
  4. Elicit (2026). Create and save columns in Elicit. Elicit Help Center. support.elicit.com/en/articles/14758162 — accessed 2026-07-27
  5. Elicit (2026). Systematic Reviews in Elicit. Elicit Help Center. support.elicit.com/en/articles/14759154 — accessed 2026-07-27
  6. SciSpace (2026). Chat with any PDF | AI-powered Research Paper & PDF Summarizer. scispace.com/chat-pdf — accessed 2026-07-27
  7. SciSpace (2026). AI-Powered Data Extraction from Research PDFs. scispace.com/extract-data — accessed 2026-07-27
  8. Paperpal (2026). AI Academic Writing Tool — Comprehensive AI Research Assistant. paperpal.com/ — accessed 2026-07-27
  9. SCiNiTO (2026). PDF Analysis. SCiNiTO user guide. scinito.ai/help/pdf-analysis — accessed 2026-07-27
  10. Zotero (2026). Plugins for Zotero. Zotero Documentation. zotero.org/support/plugins — accessed 2026-07-27
  11. Zotero (2026). Zotero PDF Reader and Note Editor. Zotero Documentation. zotero.org/support/pdf_reader — accessed 2026-07-27
  12. Davis, Lewenstein, Simon, Booth & Connolly (2008). Open access publishing, article downloads, and citations: randomised controlled trial. BMJ 337, a568. doi.org/10.1136/bmj.a568 — accessed 2026-07-27
  13. Ottaviani (2016). The Post-Embargo Open Access Citation Advantage. PLOS ONE 11(8), e0159614. doi.org/10.1371/journal.pone.0159614 — accessed 2026-07-27
  14. Piwowar et al. (2018). The state of OA: a large-scale analysis of the prevalence and impact of Open Access articles. PeerJ 6, e4375. doi.org/10.7717/peerj.4375 — accessed 2026-07-27
  15. Langham-Putrow, Bakker & Riegelman (2021). Is the open access citation advantage real? A systematic review. PLOS ONE 16(6), e0253129. doi.org/10.1371/journal.pone.0253129 — accessed 2026-07-27
  16. Huang et al. (2024). Open access research outputs receive more diverse citations. Scientometrics 129(2), 825–845. doi.org/10.1007/s11192-023-04894-0 — accessed 2026-07-27
  17. Peters, Bertazzoli, DeJesus, van der Velden & Chin-Yee (2026). Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models. Public Underst. Sci. (online 20 Apr 2026). doi.org/10.1177/09636625261425891 — accessed 2026-08-21
  18. Anthropic (2026). System Card: Claude Opus 5 (24 July 2026; vendor self-evaluation). www-cdn.anthropic.com — accessed 2026-08-21