AI for Researchers · Visual Deck
How large language models work, where they fail, and how to tell the difference
{{Presenter Name}}
Landscape as of August 2026
2-Minute Orientation
flowchart LR s1(["1 · Foundations
today"]):::today --> s2(["2 · Literature
discovery"]):::next s2 --> s3(["3 · Reading
& notes"]):::rest s3 --> s4(["4 · Writing
& integrity"]):::rest s4 --> s5(["5 · Data
& code"]):::rest s5 --> s6(["6 · Multi-
disciplinary"]):::rest s6 --> s7(["7 · Discipline
deep-dives"]):::rest classDef today fill:#ab7d22,stroke:#d9b36c,color:#14132b,font-size:19px classDef next fill:#23204c,stroke:#d9b36c,color:#eceaf8,font-size:19px classDef rest fill:#23204c,stroke:#8f8cb8,color:#b9b7d6,font-size:19px
~20 min foundation · ~25 min live demo · ~15 min Q&A — every session.
Today builds the two artefacts the whole series reuses: the capability/failure map and the 2026 tool taxonomy.
Learning Objectives
Objectives 4–6 next · full verbatim wording in the reference deck and curriculum
Learning Objectives
Full verbatim wording in the reference deck and curriculum
Section 01
What the machine is actually doing when it answers your research question — and why that explains almost every failure you will meet.
Literacy Foundation
Wiley ExplanAItions, fielded Aug 2025 (latest wave; vendor self-reported) [16] · peer-reviewed corroboration [15]
Literacy Foundation
flowchart LR P["“The study was published by …”
your prompt, as tokens"]:::prompt --> M(["statistical patterns"]):::model M -- "p = .41" --> A["Smith ▊"]:::top M -- "p = .22" --> B["Gorman ▊"]:::mid M -- "p = .04 · long tail…" --> C["Nguyen ▊"]:::low classDef prompt fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef model fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef top fill:#23204c,stroke:#d9b36c,color:#d9b36c,font-size:20px classDef mid fill:#23204c,stroke:#8f8cb8,color:#b9b7d6,font-size:20px classDef low fill:#23204c,stroke:#55527e,color:#8f8cb8,font-size:20px
Literacy Foundation
"Comparative genomics of penguins" —
Literacy Foundation
outside the window: your other chats · yesterday's upload · the rest of the literature —
it exists only if the product deliberately re-injects it (saved memory, project files)
separate limit: reliable knowledge can end months before today's date [2] [3]
Literacy Foundation
Put what matters at the top or bottom — and five well-chosen papers beat fifty dumped in. Measured on the 2023-era cohort [9]; no current-generation replication, so assume it still applies.
Literacy Foundation
flowchart TB
M(("statistical patterns
of plausible text")):::model --> R["Gorman et al. (2014)
real"]:::real
M --> F["Sørensen et al. (2019)
invented — same font, same confidence"]:::fake
classDef model fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px
classDef real fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px
classDef fake fill:#1c1a3f,stroke:#c14b58,color:#ef9a9a,font-size:19px
Confident, well-formed and wrong is the expected failure mode — not an anomaly [4]
Literacy Foundation
how most benchmarks score an uncertain model [4]
models are "optimized to be good test-takers" — guessing raises the score
What you can do today: give the model explicit permission to abstain, in your prompt [4]
Literacy Foundation
Literacy Foundation
flowchart LR P["your prompt"]:::io --> T["hidden thinking tokens
─ effort dial: minimal ⟶ max ─"]:::think --> A["answer"]:::hot classDef io fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:20px classDef think fill:#1c1a3f,stroke:#8f8cb8,color:#b9b7d6,stroke-dasharray:6 6,font-size:20px classDef hot fill:#23204c,stroke:#d9b36c,color:#d9b36c,font-size:20px
Literacy Foundation
Raise the dial for hard multi-step work; verify facts exactly as before.
Section 02
One table you can hold in your head: which research tasks are safe to hand to AI, which are unreliable, and which are simply dangerous.
Series Artefact
"Safe" means safe to draft with — never safe to skip reading. Fabrication is structural; vendors document it in their own cards [5] [7] [27]
Series Artefact
Section 03
Three tiers, one question: before you type anything, know which kind of tool you have just opened.
Series Artefact
One caveat: Tier-2 tools such as Consensus and Scite now ship MCP servers, so a Tier-1 chatbot can call them directly [30] [31] — read the tiers as what you are reaching for, not walls. Examples illustrative, not endorsements.
Series Artefact
Tier 1 — generate first
flowchart LR G["GENERATE"]:::hot --> S["search
(bolted on, maybe)"]:::ghost classDef hot fill:#23204c,stroke:#d9b36c,color:#d9b36c,font-size:20px classDef ghost fill:#1c1a3f,stroke:#55527e,color:#8f8cb8,stroke-dasharray:6 6,font-size:19px
leans on memory even when search runs [7] · its library ends at the knowledge cutoff [2] [3] · excellent at language, structure, brainstorming
Tier 2 — retrieve first
flowchart LR S2["SEARCH
scholarly corpus"]:::safe --> G2["generate"]:::norm classDef safe fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:20px classDef norm fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:20px
"We only use AI after we search the scientific literature" [19] · built on Semantic Scholar, OpenAlex [19] [21] · still summarises — still flattens nuance (Session 3)
Series Artefact
flowchart LR P["plan"]:::n --> S["search"]:::n --> R["read"]:::n --> I["iterate…"]:::n I --> S I --> REP["cited report
minutes, not seconds"]:::hot classDef n fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef hot fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px
More citations per query — and worse odds each [8]. Some let you edit the plan before it runs [23]. Treat the report as a well-organised lead list, never a verified synthesis.
Section 04
Not magic words — a repeatable structure, an iteration habit, and a routine for evaluating what comes back.
Systematic Prompting
Then iterate: a prompt is a first draft, not an incantation.
Systematic Prompting
"Summarise the literature on AI adoption by researchers."
no field · no window · no shape · nothing about being unsure — an invitation to invent references [5]
"I study scholarly communication and have attached three survey papers. Writing for a departmental seminar audience, use only the attached papers — if a figure is not in them, say so rather than estimating. Produce a 200-word synthesis with a page reference for every claim."
abstention is permitted, not penalised [4]
— Context · — Role & register · — Instructions & constraints · — Task
Systematic Prompting
Prompting is a documented literature — 58 techniques catalogued [13]. Context, constraints and retrieval buy correctness.
Systematic Prompting
flowchart LR A["1 · Grounding
where did this come from?
nothing retrieved = nothing checked"]:::n --> B["2 · Citations
does every reference resolve
to a record you can open?"]:::n --> C["3 · Claim-to-source
does the source say
what the output claims?"]:::n --> D["4 · Omission
what was left out?
summaries lose caveats first"]:::n classDef n fill:#23204c,stroke:#d9b36c,color:#eceaf8,font-size:19px
Checks 2–4: [5] [7] [6] [18] — a real link can still carry a false claim; open the record, don't trust the format.
Live Demo & Takeaways
References
References