AI for Researchers
How large language models work, where they fail, and how to tell the difference
{{Presenter Name}}
Landscape as of August 2026
2-Minute Orientation
Today (Session 1): the mental model every later session depends on — and the two artefacts we will reuse all series, the capability/failure map and the 2026 tool landscape taxonomy.
Learning Objectives
Objectives 4–6 continue on the next slide.
Learning Objectives
Section 01
What the machine is actually doing when it answers your research question — and why that explains almost every failure you will meet.
Literacy Foundation
Adoption figures: Wiley ExplanAItions survey, fielded Aug 2025 — still the latest wave as of Aug 2026; publisher self-reported data [16]. Barrier figure: Ithaka S+R, self-reported survey data [17]. Peer-reviewed corroboration: [15].
Literacy Foundation
Literacy Foundation
Literacy Foundation
Literacy Foundation
Literacy Foundation
Literacy Foundation
Literacy Foundation
Literacy Foundation
Literacy Foundation
Section 02
One table you can hold in your head: which research tasks are safe to hand to AI, which are unreliable, and which are simply dangerous.
Series Artefact
| Zone | Typical research tasks | Why it lands here | Your obligation |
|---|---|---|---|
| Safedraft with it, still read it | Polishing, translating and restructuring text you wrote; brainstorming search terms; explaining a concept you can check | You supply the ground truth, so errors are visible to you | Read every line. Disclose AI writing assistance where the journal requires it [25] |
| Unreliableneeds checking, every time | Summarising a paper you have; first-pass title/abstract screening; drafting extraction tables; generating analysis code | Evidence is mixed and task-dependent; validated applications remain rare [18] | Spot-check against the source; never let AI be the only pass [18] |
| Dangerousdo not delegate | Asking for references from memory; treating an AI synthesis as evidence; uploading confidential manuscripts or grant applications | Fabrication is structural, and vendors document it in their own model cards [5] [7] [27]; funders prohibit AI in peer review [26] | Retrieve citations from a database instead; keep confidential material out of these tools [26] |
Series Artefact
Section 03
Three tiers, one question: before you type anything, know which kind of tool you have just opened.
Series Artefact
| Tier | What it is | Examples, August 2026 | Where it fails |
|---|---|---|---|
| Tier 1General chatbots | An LLM in a chat box; may search the web if asked | ChatGPT (GPT-5.6), Claude (Fable 5 / Opus 5), Gemini (3.1 Pro / 3.7 Flash), Kimi (K3) | Leans on memory even when search is available [7]; worst for citations [5] |
| Tier 2Research-specific tools | Retrieval over a scholarly corpus first, generation second [19] | Elicit, Consensus, Semantic Scholar, SCiNiTO, Scite, ResearchRabbit, Gemini Notebook [19] [20] [21] [22] | Bounded by corpus coverage; paywalled full text often missing |
| Tier 3Deep-research agents | Plans, browses, iterates, returns a long cited report [23] | ChatGPT deep research, Gemini Deep Research [23] [24] | More citations per query and a higher URL-hallucination rate [8] |
Series Artefact
Series Artefact
Section 04
Not magic words — a repeatable structure, an iteration habit, and a routine for evaluating what comes back.
Systematic Prompting
Systematic Prompting
"Summarise the literature on AI adoption by researchers."
"I study scholarly communication and have attached three survey papers. Writing for a departmental seminar audience, use only the attached papers — if a figure is not in them, say so rather than estimating. Produce a 200-word synthesis with a page reference for every claim."
Systematic Prompting
Systematic Prompting
Live Demo & Takeaways
References
References