AI for Researchers
Grounded notebooks, structured extraction, and the nuance a summary quietly drops
{{Presenter Name}}
Landscape as of August 2026
2-Minute Recap
Today (Session 3): the shortlist becomes a corpus — and we find out what a summary of it leaves out.
Learning Objectives
Objectives 4–6 continue on the next slide.
Learning Objectives
Section 01
Before we trust an answer about a PDF, it is worth knowing which part of the PDF the tool was allowed to look at.
Literacy Foundation
Google states the retrieval step plainly: the notebook "retrieves the most relevant information based on your question first, then builds a response with this information" [15]. Step 2 is where most failures are born, and it is the step you cannot see. One 2026 update: million-token flagships (Session 1) can now ingest a small corpus whole — no retrieval step — but that removes lost passages, not misreading, and at notebook scale retrieval is still what happens.
Tier 2 — Research-Specific Tools
| Tool | What it is built to do | How it shows its evidence | Documented limit |
|---|---|---|---|
| Gemini Notebookformerly NotebookLM [14] | Source-bound Q&A over a corpus you upload [15] | In-line citations; hover for the quote, click to jump to it [16] | Free tier: 50 sources per notebook, 50 chat queries/day [18] |
| Elicit | Extraction tables: one column per data point across many papers [20] | Click any cell to see the supporting quote from the paper [21] | Without the PDF it "can only pull information from the abstract" [19] |
| SciSpace | PDF chat plus a data extractor over uploaded papers [22] | "Answers backed by citations from specific sections of the PDF" [22] | Maximum 50 extraction columns per table [23] |
| Paperpal | Chat across uploaded sources, inside a writing workflow [24] | "Get key findings with citations" [24] | Writing-workflow scope; corpus limits not documented [24] |
| SCiNiTO | Structured insights from an uploaded article, plus follow-ups [25] | Analysis history keeps every prior PDF and its insights [25] | Per-article analysis; cross-corpus use not documented [25] |
| Zotero + plugins | Your library is the corpus; AI arrives as third-party code [26] | Annotations carry page links and citations into notes [27] | "Plugins have full access to your Zotero and your computer" [26] |
Critical Literacy
Section 02
Four independent lines of evidence, all pointing at the same thing: the qualifier is the first casualty.
The Evidence — 1 of 4
Numbers above: the 2025 ten-model study [1]. No study yet tests the August-2026 cohort on overgeneralisation — an open gap, not an all-clear.
The Evidence — 2 of 4
Qualitative study, 20 interviews, reported as a refereed conference poster; summaries were ChatGPT-4 (2025). Cited for which content disappears, never as a rate.
The Evidence — 3 of 4
Standardised test passages, not journal articles. Take the direction; do not quote the effect size for research reading.
The Evidence — 4 of 4
Series Artefact
Rule of thumb: use summaries to decide what to read and what to compare — never as the thing you cite.
Section 03
A corpus you chose, a model that may only answer from it, and a citation on every sentence. What that buys — and what it does not.
Grounded Notebooks
Vendor self-reported adoption, quoted as such: Google states "more than 30 million people and over 600,000 organizations" use it [14]. Nothing on this slide depends on that number.
Critical Literacy
Grounded Notebooks
Section 04
Twenty papers, one grid: methods, populations, effect sizes. The part of AI-assisted reading with the best evidence behind it — and the strictest conditions.
Structured Extraction
This is the same discipline as a systematic-review extraction form. The tool changes; the form does not.
The Evidence
Series Artefact
Name: one phrase — "Primary outcome measured in the study".
Instruction: where to look, what counts, what format.
Precision: "Be as precise as possible — '50 hours' is better than '2 days'." [20]
Abstain rule: "If the paper does not state this, write NOT REPORTED. Do not infer it."
Today's Demo Corpus
| Paper | Design | Population / sample | Headline result |
|---|---|---|---|
| Davis et al. 2008BMJ 337, a568 | Randomised controlled trial | 1,619 articles, 11 physiology journals | +89% full-text downloads; no citation advantage at one year [28] |
| Ottaviani 2016PLOS ONE 11(8) | Observational, matched control | 3,850 opened vs 89,895 closed articles | Advantage "as high as 19%"; bigger for above-median articles [29] |
| Piwowar et al. 2018PeerJ 6, e4375 | Large-scale observational | 3 samples × 100,000 articles | OA articles receive 18% more citations than average [30] |
| Langham-Putrow et al. 2021PLOS ONE 16(6) | Systematic review | 134 studies from 5,744 screened | 47.8% found an advantage, 27.6% found none, 23.9% in subsets [31] |
| Huang et al. 2024Scientometrics 129 | Large-scale observational | Global outputs, 2010–2019 | Different outcome: OA is cited from more places [32] |
Critical Literacy
Section 05
Notebooks come and go. Your reference manager is where the reading has to land if you want it in two years' time.
Knowledge Management
Systematic Prompting
Name the section, ask one thing, and demand the quote.
Ask for a grid, per source, with disagreements kept intact.
Live Demo & Takeaways
References
References