← series index
# Demo Script — Session 2: Literature Discovery & Synthesis
*AI for Researchers — presenter script · Landscape as of August 2026*
**Runtime:** ~25 minutes (Demo A ~7 min · Demo B ~8 min · Verification ~10 min)
**Slides it follows:** slide 28 ("What We Do Next")
**Artefacts it exercises:** The Literature-Discovery Tool Matrix (slides 13–14) and The Citation-Verification Workflow (slide 26)
---
## 0. The research question
Everything in this demo runs on **one** question, chosen deliberately:
> **Can AI-based literature search tools replace traditional Boolean database searching for systematic reviews?**
Why this question:
- It is **methodological**, so it lands for any discipline in the room.
- It is **genuinely contested**, so a deep-research agent has something to be one-sided about.
- **You already hold the ground truth.** The answer is sources [2], [3], [27] and [28] in this session's `sources.md`. When a tool returns something, you can check it live against papers you are holding, rather than hoping.
- The audience has just seen the numbers on slides 9 and 25, so they can score the tools themselves.
> **Say this out loud before you start:** "I am about to ask two tools a question I already know the answer to. That is the only honest way to demo a search tool."
---
## 1. Preparation (do this the day before — not the morning of)
| # | Item | Why |
|---|---|---|
| 1 | Log in to your chosen discovery tool and your chosen agent; confirm both work on the venue network. | Institutional SSO and conference wifi are the two most common live failures. |
| 2 | Run **both** demos start to finish and save the outputs. | These become your fallback (§6). |
| 3 | Note how long the agent took. | You will need to fill that time on the day — see §4, step 4. |
| 4 | Pre-resolve the three citations you will verify in §5. | So you know in advance which one is wrong, and why. |
| 5 | Open your institutional-access proxy / library link resolver in a spare tab. | §5 step 3 needs full text, and a paywall wall on screen kills momentum. |
| 6 | Have `sources.md` open on your second screen. | It is the answer key. |
| 7 | Set browser zoom so 20px body text is readable from the back of the room. | Both tools render small by default. |
**Tool neutrality note.** The script below names specific tools because a script has to. Say so on the day: *"I am using X because I have an account; every tool on slide 13 does this job and you should pick on the basis of your corpus and your budget, not my demo."* Any Tier-2 tool from slide 13 substitutes into Demo A; any of the three agents on slide 17 substitutes into Demo B.
---
## 2. Demo A — A citation-grounded discovery tool
**Tool used in this script:** Consensus (free tier, no institutional login needed).
**Substitutes:** Elicit, SCiNiTO, Scite Assistant, Semantic Scholar + Scite for the citation-context step.
### Step A1 — The naive query (30 seconds)
Type exactly this into the search box:
```
AI literature search
```
**Expected outcome:** a broad, unfocused result set — plenty of papers, none of which answer your actual question. Do not dwell. The point is one sentence: *"This is what most people type, and it is why they conclude the tool is useless."*
### Step A2 — The real query
Clear the box. Type exactly this:
```
Can AI-based literature search tools achieve the same recall as traditional Boolean database searching in systematic reviews?
```
**Expected outcome:**
- A synthesised answer paragraph with numbered inline citations.
- A ranked paper list beneath it, each with title, journal, year and a one-line finding.
- Consensus should surface evaluation studies of AI search tools; the Lau & Golder paper [3] is the single most on-point result and often appears in the top ten.
**What to point at on screen:**
1. **The citations came first.** Read the slide-12 line back to them: the tool retrieves before it generates [1].
2. **The corpus boundary.** Whatever is not in the 200M-document corpus cannot appear, no matter how well you phrase the question [1].
3. **The ranking is a judgement.** Consensus's 2025 pipeline re-ranked on recency, citation count and journal impact before you saw anything [1]; the 2026 Scholar Agent rebuild [35] changes the machinery, not the fact that something ranks first. A 2016 methods paper can be ranked below a 2025 one that says less.
### Step A3 — Constrain with the interface, not the prose
Apply the tool's own filters: **study type**, **year range**, **open access**.
**Say:** *"This is the 'I' in CRIT. In a discovery tool you constrain with the interface, because the filters are deterministic and the prose is not."*
### Step A4 — The honest scorecard
Before moving on, fill this in live:
| Question | Answer on the day |
|---|---|
| Did it find Lau & Golder (2025)? | |
| Did it find Bramer et al. (2016)? | |
| Did it find the Cochrane/Campbell/JBI/CEE position statement? | |
| How many of the top 10 were actually about *evaluating* AI search, rather than *using* it? | |
**Expected outcome — and the teaching point (dual path; you cannot know in advance which you get).**
| If the scorecard shows… | Say |
|---|---|
| Some hits, some misses | "That is the 39.5%-sensitivity picture from slide 9 [3] — measured in early 2025 — reproduced live. Ranking is not enumerating." |
| All three found | "Today it found everything I planted. That does not overturn slide 9; it confirms slide 20 — the same tool scored 25.5% and 69.2% depending on the question [3]. Four case studies and one demo are both small samples. The scorecard, not the score, is the lesson." |
Run the check either way; the habit you are modelling is the check itself.
---
## 3. Interlude — hand the same question to the agent (1 minute)
**Say:** *"Demo A took ninety seconds and gave me a shortlist. Now I am going to spend eight minutes and get an essay. Watch what I gain, and watch what I give up."*
---
## 4. Demo B — A deep-research agent
**Tool used in this script:** ChatGPT deep research — which runs a GPT-5.2-based model, not the GPT-5.6 chat flagship; worth saying out loud [21][34].
**Substitutes:** Gemini Deep Research or Deep Research Max (both on Gemini 3.1 Pro) [22][32], Claude Research (paid plans; model not publicly named) [23]. All three plan-then-browse-then-report [21][22][23].
### Step B1 — The prompt (verbatim)
```
Context: I am preparing a methods seminar for academic researchers across
disciplines. The question is whether AI-based literature search tools
(semantic search, LLM-assisted discovery, deep-research agents) can replace
traditional Boolean database searching for systematic reviews.
Role and register: write for research-active academics who are not
information scientists. Neutral academic prose.
Instructions and constraints:
- Prioritise peer-reviewed evaluations and the published guidance of
evidence-synthesis organisations over vendor material.
- Every quantitative claim must be followed by a source with a working link.
- Present the strongest case FOR replacement and the strongest case AGAINST,
and say which is better evidenced.
- List separately, at the end: (a) any source you cited but could not read in
full, and (b) anything you looked for and could not access.
- Do not include vendor blog posts as evidence of tool performance. If you
cite one, label it as a vendor claim.
- 700 words maximum, excluding the source list.
Task: answer the question above.
```
**Why the prompt is shaped this way — say it out loud:**
- The "strongest case FOR / strongest case AGAINST" clause is a direct response to the DeepTRACE finding (a 2025 audit; no newer re-measurement of balance exists) that these systems stayed "highly one-sided on debate queries" [24]. **Assume they will not volunteer the other side. Demand it.**
- The "could not read in full / could not access" list is the paywall probe from slide 24. Roughly four in five catalogued works are not open access [20], and an agent will rarely tell you that unprompted.
- The vendor-material clause applies to the agent the same rule this deck applies to itself: no vendor's self-reported performance claim is treated as a measurement. Every performance figure on slides 9, 19 and 20 comes from independent evaluation [3][24][25][31].
### Step B2 — Edit the plan before it runs
The agent proposes a research plan. **Do not accept it.** Read it aloud and change at least one thing — narrow a sub-question, or add "include evidence-synthesis methodology guidance".
**Say:** *"This is the highest-leverage thirty seconds in the whole workflow, and almost nobody uses it"* [21][22].
### Step B3 — While it runs (this is not dead air)
It will take several minutes [22]. Use the time:
- Take Q&A on Demo A.
- Or walk back through slide 19 and tell the room what you are about to look for.
- Or show the agent's live activity trace and narrate the searches it is choosing.
### Step B4 — Read the report critically, on screen
Work through, in this order:
1. **Length vs. substance.** Note how much of it you already knew.
2. **Balance.** Did it actually produce both cases, or did it produce one case and a hedging paragraph? [24]
3. **Density of citations.** Agents cite more per query than search-augmented chatbots — and hallucinate URLs at double the rate (10.7% pooled vs 4.8%) [25]. More retrieval is not more accuracy: fact-check accuracy dropped ~42% as tool calls scaled from 2 to 150 [31].
4. **The two lists you demanded.** Did it produce the "could not access" list? If it silently dropped that instruction, say so: *"It ignored an explicit instruction. Note that for later."*
---
## 5. The verification segment (the heart of the session — 10 minutes)
**Set-up line:** *"We are going to take three citations out of this report and run the workflow from slide 26 on each one. I do not know in advance which will fail. That is the point."*
Pick three deliberately:
- **Citation 1:** a quantitative claim (a percentage, a sample size).
- **Citation 2:** something that sounds like the paper you already hold — i.e. Lau & Golder, Bramer, or the Cochrane statement.
- **Citation 3:** the oldest or most obscure-looking reference in the list.
For each one, run all five steps on screen:
| Step | What you do | What you say |
|---|---|---|
| 1. **Resolve** | Click the link. | "In 2025 audits, 5–18% of these did not resolve [25]. Current frontier systems keep link validity above 94% [31] — so if this opens, we have learned almost nothing yet. Resolving is necessary, not sufficient." |
| 2. **Match** | Compare author, year, journal, title against the record — all four. | "A real paper attached to the wrong claim still looks perfect at a glance." |
| 3. **Open** | Get the full text via institutional access. Find the sentence. | "This is where the paywall bites — four of five works are not open access" [20]. |
| 4. **Compare** | Read the source sentence aloud, then the agent's sentence. | "Does it say that, or something weaker, narrower or opposite? This is where current systems fail: with valid links, factual support runs 39–77%" [24][31]. |
| 5. **Direction** | Check whether later work supported or contradicted it (citation-context tool). | "Real, correctly quoted, and superseded is still wrong" [13]. |
**Expected outcomes, in rough order of likelihood (August-2026 base rates — run the checks; either branch teaches):**
| Outcome | How likely | What to say |
|---|---|---|
| Citation is real and fairly represented | Most common per citation | "Good. That is the base rate we want — and across 14 current models, factual support still only runs 39–77%" [31]. |
| Citation is real but the claim is stronger than the source supports | **The most likely failure** — plan your commentary around it | "This is the failure citation-grounding does **not** fix [1] — and on the current generation it is the dominant one: the link works, the paper is real, the claim is not supported" [31]. Point back at slides 12 and 19. |
| A number in the report does not appear anywhere in the cited paper | Occasional | The single most dangerous outcome, because it is invisible unless you open the paper. Same family as the row above. |
| Link is dead or redirects to a landing page | Now uncommon on frontier agents (link validity >94% [31]; 5–18% in 2025 audits [25]) | Distinguish link rot from fabrication — check the archive. |
| Citation does not exist at all | Rare on current frontier systems — do not promise it | "3–13% of citation URLs were hallucinated in 2025-era audits [25]. If we see one today it is a bonus, not the lesson — the lesson is the row two up." |
If all three citations pass all five steps, that is the *expected-good* branch, not a failed demo: say the pass rate out loud, then show the pre-captured rehearsal failure (§6) so the room sees both branches.
### Closing the loop (2 minutes)
Put the two outputs side by side and ask the room:
> **"Which of these would you cite in a grant application tomorrow?"**
Land it on:
1. The discovery tool gave a **shortlist you can check**. The agent gave **prose you must check**.
2. Neither is a search strategy. Cochrane, Campbell, JBI and CEE all require human oversight and transparent reporting of any AI use that makes or suggests a judgement [27].
3. You verified three citations in ten minutes. A 40-reference report is a two-hour job — **budget it, or do not cite it.**
---
## 6. Fallback plan (assume something breaks)
| If this fails | Do this instead |
|---|---|
| Venue network is down or slow | Run entirely from the pre-captured screenshots (below). Narrate as if live; say you are using captures. |
| Discovery tool requires a login you cannot complete | Switch to Semantic Scholar — no account needed [12] — and do the citation-context step with the free Scite badge if available. |
| Agent is queued, rate-limited, or exceeds your usage allowance | Open the pre-run report from the day before. Usage limits are quota-based and plan-dependent — OpenAI's last published per-tier figures date to mid-2025, so check your own account beforehand, not a slide [21][34]. |
| Agent takes far longer than rehearsal | Cut §4 step B4 to items 2 and 4, and protect the verification segment. **Never cut §5.** |
| Every citation checks out clean | Say so, honestly — then verify a fourth, and if that is clean too, show the pre-captured failure from rehearsal and note that rates are per-query, not per-tool. |
| A tool has changed its interface since rehearsal | Say "this changed since last week" and move on. It reinforces the freshness caveat on every slide. |
| Someone asks about a tool not in the matrix | "It is not in the matrix because I have not read its documentation, not because it is bad." Add it to the follow-up list. |
### Pre-capture screenshot checklist (capture all of these the day before)
- [ ] Demo A: the naive query result set
- [ ] Demo A: the real query — synthesised answer with inline citations visible
- [ ] Demo A: the paper list, showing at least one hit and one miss against your answer key
- [ ] Demo A: filters applied (study type / year / open access)
- [ ] Demo B: the agent's **proposed plan**, before editing
- [ ] Demo B: the plan **after** your edit
- [ ] Demo B: the activity trace mid-run
- [ ] Demo B: the finished report, top of page
- [ ] Demo B: the report's source list, in full
- [ ] Demo B: the "could not access" section — or the visible absence of it
- [ ] Verification: one citation resolving cleanly, with the supporting sentence highlighted in the source
- [ ] Verification: one citation failing (dead link, mismatched metadata, or unsupported claim) — **this is the one you must have**
- [ ] Verification: a citation-context view showing supporting vs. contrasting citations
- [ ] A paywall interstitial, to make the slide-24 point concretely
---
## 7. Presenter notes
**On tool neutrality.** Six tools are on the matrix and you will demo two. Say early that the choice is about your account, not their quality, and repeat it once at the end. If someone asks "which is best", the honest answer is on slide 14: it depends on whether you need extraction tables, citation context, graph exploration or open-catalogue breadth — the strongest independent evaluation of a matrix tool [3] tested one of them, on four reviews, in early 2025, on a product since replaced [33]; the vendors' own newer benchmarks are self-graded and appear nowhere in this deck.
**On the retracted paper.** While building this session, the most on-topic-looking search result for "AI literature tool evaluation" was a **retracted** journal article. If it fits the room, say so: it is the perfect one-sentence argument for why step 5 of the verification workflow exists.
**Two objections to expect.**
1. *"This will all be fixed in six months."* Give the honest 2026 answer: it partly *was* — and watch which part. URL resolution is largely fixed: frontier systems keep link validity above 94% [31], and automated URL checking cuts non-resolving citations to under 1% in experiments [25]. What was not fixed is support: the same audited systems ran 39–77% factual accuracy, and accuracy *fell* ~42% as tool calls scaled up [31]. "Does the link work" was an engineering problem; "does the source say that" is not. The workflow you are teaching is the thing that survives model updates; the numbers are not.
2. *"Cochrane says we can use AI now, so what is the problem?"* Read key message 3 back to them verbatim: you may use it "as long as you can demonstrate that it will not compromise the methodological rigour or integrity" of the synthesis [27]. The burden of demonstration is on the author.
**What not to do.** Do not pick a research question from your own current project. You will defend the tool's answer instead of testing it, and the room will notice.
---
*AI for Researchers · Session 2: Literature Discovery & Synthesis · Landscape as of August 2026*