AI for Researchers · Visual Deck
Grounded notebooks, structured extraction, and the nuance a summary quietly drops
{{Presenter Name}}
Landscape as of August 2026
2-Minute Recap
flowchart LR s1(["1 · Foundations
the map · the taxonomy · CRIT"]):::done --> s2(["2 · Literature
the discovery matrix"]):::done s2 --> s3(["3 · Reading & notes
today"]):::today s3 --> s4(["4 · Writing
& integrity"]):::rest s4 --> s5(["5 · Data
& code"]):::rest s5 --> s6(["6 · Multi-
disciplinary"]):::rest s6 --> s7(["7 · Deep-
dives"]):::rest classDef done fill:#23204c,stroke:#55527e,color:#b9b7d6,font-size:18px classDef today fill:#ab7d22,stroke:#d9b36c,color:#14132b,font-size:19px classDef rest fill:#23204c,stroke:#8f8cb8,color:#8f8cb8,font-size:18px
Session 2's hard-won lesson: grounded ≠ correct — retrieval kills invented sources, not misread ones [11].
You now have a shortlist of papers. Today the shortlist becomes a corpus — and we find out what a summary of it leaves out.
Learning Objectives
Objectives 4–6 next · full verbatim wording in the reference deck and curriculum
Learning Objectives
Full verbatim wording in the reference deck and curriculum
Section 01
Before we trust an answer about a PDF, it is worth knowing which part of the PDF the tool was allowed to look at.
Literacy Foundation
flowchart LR P["your PDF"]:::io --> A["1 · Index
broken into passages
the tool can look up"]:::n A --> B["2 · Retrieve
the passages that
look most relevant"]:::hot B --> C["3 · Generate
an answer from
those passages"]:::n C --> D["4 · Cite
the quoted passage —
click back to it"]:::n P -.-> L["2026 second path: million-token flagships
ingest a small corpus whole — no retrieval step"]:::ghost L -.-> C classDef io fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef n fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:18px classDef hot fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:18px classDef ghost fill:#1c1a3f,stroke:#55527e,color:#8f8cb8,stroke-dasharray:6 6,font-size:16px
Google states it plainly: the notebook "retrieves the most relevant information based on your question first" [15] [16]. Step 2 is where most failures are born — and it is the step you cannot see.
The second path removes lost passages, not misreading — and at notebook scale, retrieval is still what happens.
Tier 2 — Research-Specific Tools
Every cell from that tool's own current documentation — no cell is a comparative judgement. Row order is not a ranking; no endorsement implied.
Critical Literacy
So the first question about any extraction table is: did the tool have the full text?
Section 02
Four independent lines of evidence, all pointing at the same thing: the qualifier is the first casualty.
The Evidence — 1 of 4
4,900 summaries · 10 models of the 2025 cohort (ChatGPT-4o, Claude 3.7 Sonnet, DeepSeek, LLaMA 3.3 70B) [1] — commonest failure: "the patients in this trial" becomes "patients".
Within that cohort, "newer models tended to perform worse"; the 2026 follow-up finds the bias persists [33]. No study yet tests the August-2026 cohort — an open gap, not an all-clear.
The Evidence — 2 of 4
Twenty authors compared an AI summary of their own paper with their own [2] — qualitative, ChatGPT-4 (2025); cited for which content disappears, never as a rate.
The Evidence — 3 of 4
Pre-registered randomised cross-over, n = 195, four GPT-4-era reading tools (2025) [3] — the authors: "much of the detail and nuance of the passage was lost in the summary".
Standardised test passages, not journal articles: take the direction; do not quote the effect size for research reading.
The Evidence — 4 of 4
Scientists evaluating AI answers about papers they had written themselves [5]. A structured checklist beat unaided expert reading — that is why you are getting one.
Series Artefact
Rule of thumb: use summaries to decide what to read and what to compare — never as the thing you cite.
Section 03
A corpus you chose, a model that may only answer from it, and a citation on every sentence. What that buys — and what it does not.
Grounded Notebooks
The trade you are making: a smaller world, in exchange for being able to check every sentence. Vendor self-reported adoption, quoted as such: "more than 30 million people and over 600,000 organizations" [14] — nothing here depends on that number.
Critical Literacy
RefusalBench: over 30 models of the 2024–25 cohort, six categories of flawed context, three intensity levels — "neither scale nor extended reasoning improves performance"; no re-test of the current generation yet [6]. Chat-with-paper benchmarks now score grounded refusal deliberately [7].
So: test your notebook's refusal before you trust its answers. We do exactly that in the demo.
Grounded Notebooks
Two quiet gaps: copy-protected PDFs will not import, and footnotes and comments are not imported from Google files [17] [18]
Section 04
Twenty papers, one grid: methods, populations, effect sizes. The part of AI-assisted reading with the best evidence behind it — and the strictest conditions.
Structured Extraction
flowchart LR A["define columns
the fields you would
fill in by hand anyway"]:::n --> B["pilot
on three or four
papers first"]:::n B --> C["run
the same question of
every paper — rows are papers"]:::n C --> D["audit
click a cell,
read the supporting quote"]:::hot D --> E["export
CSV · Excel · RIS
into your own analysis"]:::n classDef n fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:18px classDef hot fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:18px
A summary is one blob of prose you cannot audit; a column is a defined question asked of every paper the same way [21] [23].
The same discipline as a systematic-review extraction form — the tool changes; the form does not.
The Evidence
Six ongoing systematic reviews, blinded adjudicators, Claude 2.1→3.5 — a 2023–24 cohort — and a human checked every cell [8]. Take the human out: adjusted performance only "moderate" (GPT-4, 2024) [9].
Extraction stays Unreliable — the tradeoff persists: Opus 5 card (July 2026) reports accuracy up 11% and hallucinations up 6% [12] [34]; Cochrane, Campbell, JBI & CEE require human oversight [13]
Series Artefact
Constrain the answer space where you can — Yes / No / Maybe, or a fixed list [20]. The abstain rule is the load-bearing line: without it, a blank becomes an inference.
Today's Demo Corpus
All five open access with resolvable DOIs — re-verified against Crossref 2026-09-02; OA status via Unpaywall 2026-07-27. Full list in this session's demo script.
Critical Literacy
The review's sharpest finding is in its appraisal — of the three studies at low risk of bias: one advantage, one none, one in subsets [31].
A comparison table makes disagreement visible. A summary makes it disappear. [1]
Section 05
Notebooks come and go. Your reference manager is where the reading has to land if you want it in two years' time.
Knowledge Management
this is quote-to-source grounding — done by hand, and fully auditable
lower-risk path: export the PDFs out to the reading tool instead [19]
Systematic Prompting
Name the section, ask one thing, and demand the quote.
Ask for a grid, per source, with disagreements kept intact.
Live Demo & Takeaways
References
References