← series index
# Demo Script — Session 6: AI for Multidisciplinary Research
*AI for Researchers · presenter script · Landscape as of August 2026*
**Runtime:** ~25 minutes of the 60-minute session (following ~20 minutes of literacy foundation, leaving ~15 for Q&A).
**Artefacts it exercises:** The Cross-Field Orientation Ladder (slide 25), the false-friend table (slide 11), the Cross-Field Discovery Toolkit (slides 17–18), and the Cross-Field Evidence Compatibility Check (slide 22).
**Canvas link:** everything here fills **row 6** of Session 5's Personal AI Workflow Canvas.
---
## 0. What this demo is, and what it deliberately is not
This is **one researcher getting oriented in one unfamiliar field, in twenty-five minutes, on camera** — all four rungs of the Ladder, in order, including the checks. It is not a tour of tools. Three tools appear; each earns its place by doing something the others cannot.
**Say this out loud before you start:**
> *"I am not going to become a machine-learning researcher in twenty-five minutes, and neither are you. The goal is narrower and much more useful: to get far enough into someone else's field that I can read their papers without being fooled, and ask my new collaborator a question that is worth their time. That is the whole job."*
**The field pair.** Field A = **clinical epidemiology / prediction research**. Field B = **machine learning for clinical prediction**.
Say why you chose it: *"These two fields study the same object — a model that predicts an outcome for a patient — and they have been doing it in parallel for thirty years with different words. That makes the failure mode visible instead of theoretical. Substitute your own pair; the four rungs do not change."*
---
## 1. Preparation — the day before, not the morning of
| # | Do this | Why |
|---|---|---|
| 1 | Run all four rungs end to end yourself and **save every output**. These become the fallbacks in §7. | You will need them. Rung 2 depends on three live services. |
| 2 | Open the eight corpus papers (§2) in eight tabs, in order. | §5 requires jumping into a PDF in two seconds. |
| 3 | Have a **Connected Papers** graph and an **Inciteful Literature Connector** result pre-built and screenshotted. | Connected Papers' free tier is **5 graphs per month** `[43]`. Do not burn one live if you can avoid it. |
| 4 | Check that `incitefulmed.com/academic/` still resolves. | `inciteful.xyz` now 301-redirects there `[44]`; the old link in your notes will look broken on screen. |
| 5 | Copy the four prompts below into a scratch file. **Do not retype them live.** | Every prompt here is load-bearing; a typo changes what you are demonstrating. |
| 6 | Decide, and write down, **which two of the eight papers you will actually open** in §5. | The demo's credibility rests on this step, so do not improvise it. |
---
## 2. The corpus — eight real papers, one question
**The question the researcher brings:** *"A machine-learning model has been proposed for risk prediction in my clinical area. Should I believe it, and what would I need to see before I did?"*
All eight DOIs were verified against Crossref on **2026-07-29**. Seven are open at the publisher; the eighth has a free accepted version in an institutional repository.
| # | Paper | DOI | Why it is in the corpus |
|---|---|---|---|
| 1 | Collins, G. S., et al. (2024). *TRIPOD+AI statement*. BMJ 385:e078378. | `10.1136/bmj-2023-078378` | The reporting standard **and** the best glossary. It is where the "validation" false friend is settled in writing. |
| 2 | Sung, J., & Hopper, J. L. (2023). *Co-evolution of epidemiology and artificial intelligence*. Int J Epidemiol 52(4):969–973. | `10.1093/ije/dyad089` | An epidemiology journal publishing an epi↔AI translation table. This is a rung-1 glossary that already exists. |
| 3 | Christodoulou, E., et al. (2019). *…no performance benefit of machine learning over logistic regression…*. J Clin Epidemiol 110:12–22. | `10.1016/j.jclinepi.2019.02.004` | The headline finding of the field, and a lesson in how to state it precisely. |
| 4 | Van Calster, B., et al. (2019). *Calibration: the Achilles heel of predictive analytics*. BMC Medicine 17:230. | `10.1186/s12916-019-1466-7` | The concept field B routinely omits — a *missing* word, not a mistranslated one. |
| 5 | Wynants, L., et al. (2020). *Prediction models for diagnosis and prognosis of covid-19*. BMJ 369:m1328. | `10.1136/bmj.m1328` | What field A's appraisal machinery does when pointed at field B's output. |
| 6 | Obermeyer, Z., et al. (2019). *Dissecting racial bias in an algorithm…*. Science 366(6464):447–453. | `10.1126/science.aax2342` | The "discrimination"/"bias" collision, with consequences. Free full text via UC eScholarship. |
| 7 | Paulus, J. K., & Kent, D. M. (2020). *Predictably unequal…*. npj Digital Medicine 3:99. | `10.1038/s41746-020-0304-9` | The paper that names the collision explicitly. Pairs with #6. |
| 8 | Kapoor, S., & Narayanan, A. (2023). *Leakage and the reproducibility crisis in machine-learning-based science*. Patterns 4(9):100804. | `10.1016/j.patter.2023.100804` | The cross-field failure mode itself — and, usefully, a paper about how terminology obstructed its own review. |
**Two spares**, if the room asks for more: Riley, R. D., et al. (2020), *Calculating the sample size required for developing a clinical prediction model*, BMJ 368:m441 (`10.1136/bmj.m441`); and Collins, G. S., et al. (2025), *Clinical prediction models using machine learning in oncology*, BMJ Oncology 4(1):e000914 (`10.1136/bmjonc-2025-000914`).
---
## 3. Rung 1 — Translate (~5 minutes)
### Prompt 1 (paste verbatim)
```
Context: I am a clinical epidemiologist with no training in machine
learning. I need to read machine-learning papers on clinical prediction
models well enough to judge whether their evidence bears on my question.
Role and register: Act as a methodologist who works in both fields,
explaining to a competent researcher from the other one. Assume strong
statistical literacy. Assume no machine-learning knowledge at all.
Instructions and constraints:
- Give me the 20 terms I cannot read a machine-learning clinical-
prediction abstract without.
- For each: one sentence, plain language.
- Then, separately, FLAG EVERY TERM THAT ALSO EXISTS IN CLINICAL
EPIDEMIOLOGY WITH A DIFFERENT MEANING, and give both meanings side by
side in a two-column table.
- Also list terms where the two fields use DIFFERENT WORDS FOR THE SAME
THING.
- Mark any entry you are less than confident about with "LOW CONFIDENCE".
- Do not cite any paper, and do not name any author. I will source these
myself.
- Do not use web search for this step.
Task: Produce the glossary, then the false-friend table, then the
synonym table.
```
### Expected outcome
You should get a serviceable glossary and a false-friend table that, in rehearsal, reliably **includes** *bias* and *validation*. Whether it catches or softens *discrimination* and *calibration* is **run-dependent on current models** — some runs catch all four. Whatever it produces, say the same thing:
> *"That is a good draft. It is worth nothing until I check it."*
If it caught all four false friends, the check below still earns its place — say: *"I only know this table is right because I am about to check it against the field's own guideline. Last run it read just as well; I had no way to tell the runs apart until I checked."*
### The check — do this on screen, it is the point of the rung
Open corpus paper **#1 (TRIPOD+AI)** and read Box 1 aloud:
> *"There is no such thing as a validated prediction model. **To avoid ambiguity and harmonise terminology, we refer to validation as evaluation in this article.**"* `[57]`
And the Box 1 footnote:
> *"**Validation data often has different meanings.** For example, in machine learning studies, validation data can refer to **data used for parameter tuning** or data used to evaluate model performance."* `[57]`
**Land it:** *"An international reporting guideline with 34 authors had to rename a word because two fields could not agree on it. That is not a communication problem I am going to solve with a better prompt. It is a fact about the literature, and it is exactly what a glossary is for."*
Then open **#2 (Sung & Hopper)** and put its Table 1 next to the AI's table. Note the rows the model missed — *internal validation ↔ validation*, *external validation ↔ test*, *risk factors ↔ features*, *winner's curse ↔ overfitting* — and its bias sentence:
> *"In epidemiology, bias means **systematic errors** … **In AI, bias means the intercepts in unit models.**"* `[63]`
**Say:** *"Three meanings for 'bias' before we have even reached the fairness literature. The glossary is the deliverable of rung 1. Not the summary — the glossary."*
### The trap to name out loud
If your glossary reads beautifully, that is not evidence it is right. In a study of 150 readers (pre-2026 models), LLM-written plain-language summaries were rated **as clear as human-written ones** while producing **significantly worse comprehension** `[23]`; across 4,900 summaries from the 2024–25 cohort, models overgeneralised in **26–73%** of cases, and **asking for accuracy roughly doubled it** `[22]` — a bias the same team re-confirmed on GPT-5-class models in 2026 `[64]`.
---
## 4. Rung 2 — Map (~6 minutes)
### Step 2a — the similarity graph
Seed **Connected Papers** with corpus paper **#3 (Christodoulou)** — a paper you already trust.
**What to point at on screen:** the clusters, and then this sentence from the tool's own About page:
> *"**Connected Papers is not a citation tree.**… Our similarity metric is based on the concepts of **Co-citation and Bibliographic Coupling**."* `[43]`
**Say:** *"Two papers can sit next to each other here having never cited each other. That is a feature — it is how you find the corner of field B that is working on your problem under another name."*
### Step 2b — the connector (the cross-field move)
Open **Inciteful's Literature Connector** and enter **two** papers, one from each side: **#3 (Christodoulou, field A)** and **#8 (Kapoor & Narayanan, field B)**.
Read the tool's own framing off the page: *"**Interested in interdisciplinary studies?** Discover how two bodies of literature connect to one another through citations."* `[44]`
**What to point at:** the intermediate papers on the shortest path. Those are the bridge literature — the people already doing the translation. Note the documented limits while it renders: shortest paths only, one paper per end, at most six hops `[44]`.
### Step 2c — the honest scorecard (60 seconds, do not skip)
Say the numbers:
> *"A preprint — so read it as an indication, not a settled figure — ran the same query twice in five of these tools and got back **11.8% to 28%** the same papers, at **21% to 41%** precision." `[28]` "You do not need the exact number to act on it. One run is a sample, not a search. That is why rung 2 says two tools, and why they have to be structurally different: a similarity graph and a shortest-citation-path fail in different directions, and two semantic search boxes fail in the same direction."*
If someone asks why not just ask a chatbot: a preprint testing 27 models (through the 2025 generation), 155 topics and 200 real prompt templates found **all models less epistemically diverse than a basic web search** `[27]` — no equivalent test of the August-2026 cohort exists yet.
---
## 5. Rung 3 — Read (~7 minutes)
### Prompt 2 (paste verbatim)
```
Context: I am a clinical epidemiologist orienting into machine learning
for clinical prediction models. I have a landscape map and a glossary.
Instructions and constraints:
- Propose a reading list of exactly 8 papers that would orient me.
- Rank them in the order I should read them.
- For each, give: (a) the full citation, (b) a DOI, (c) two sentences on
why it is on the list, (d) one sentence on what I can safely skip, and
(e) whether it is open access.
- Prefer reporting guidelines, systematic reviews and methodological
critiques over individual model-development papers.
- If you are not certain a DOI is correct, write "DOI UNVERIFIED"
instead of guessing.
- Prefer papers at least one year old. I want the field's established
canon, not this month's preprints.
Task: Produce the ranked, annotated reading list as a table.
```
*(That last constraint used to read "nothing from the last 6 months" — a guard against cutoff-era blind spots, where a model would fabricate what it could not have seen. Live-search assistants no longer have that blind spot; the line now guards against the opposite failure — a canon assembled from whatever surfaced this week.)*
### The check — this is the segment the room remembers
**5a — Resolve every DOI, live (~3 min).** Paste each DOI into `doi.org` in front of the room.
**Expect them all to resolve.** On a current live-search assistant that is the likely outcome — and it is the teaching moment, not the anticlimax:
> *"Every link works. Two years ago that would have ended the check. Today it barely starts it: a May-2026 audit of 14 current models found even the strongest kept link validity above 94% — while factual support for the claims those links carried ran 39% to 77%* `[66]`*. The test has moved from 'does the paper exist' to 'does the paper say what the annotation says'. That is 5b, and it is why 5b can never be cut."*
**If a DOI does miss, or resolves to the wrong paper** — better still; slow down and show it. Then give the gradient: in 2025, fabrication in AI-generated literature reviews ran **6% for a densely published topic and 28–29% for two niche ones**, same model, same task `[21]`. **Field B is your niche topic** — rarity still moves the risk, whatever the year's absolute rates.
**5b — Read two of the eight against their annotation (~4 min).** Open the two you chose yesterday. Suggested pair, because both have a checkable claim in the abstract:
- **#3 Christodoulou.** The annotation will say "found no benefit of ML over logistic regression". Open it and read the precise version: *"We found no evidence of superior performance of ML over LR"* — but at **low** risk of bias the difference in logit(AUC) was **0.00 (95% CI −0.18 to 0.18)**, while at **high** risk of bias it was **0.34 (0.20–0.47) higher for ML** `[60]`. **Say:** *"The annotation was true and useless. The finding is not 'ML doesn't work' — it is 'the apparent advantage lives in the studies at high risk of bias'. That distinction is the entire paper, and it did not survive summarisation."*
- **#4 Van Calster.** Check whether the annotation mentions **calibration at all**. Then give the number from #3: calibration *"was not addressed in 56 (79%)"* of the 71 studies reviewed `[60]`.
**The rule to state:** *"If either annotation misdescribes its paper, I discard the whole list and start rung 3 again. Not because the tool is bad — because I now have no way of knowing which of the other six are wrong."*
---
## 6. Rung 4 — Synthesise (~5 minutes)
### Prompt 3 (paste verbatim) — note what it is *not* given
```
Context: Below are MY OWN notes on 8 papers I have read about machine
learning for clinical prediction models. I am a clinical epidemiologist.
[PASTE YOUR NOTES]
Instructions and constraints:
- Write a one-page memo for a colleague in MY field, not in the other one.
- Structure it as: (1) what this field claims; (2) what its own critics
say; (3) the three terms where our fields mean different things; (4)
what I would need to see before believing a model from this field.
- Every claim must be attributable to one of MY notes. If a claim is not
in my notes, do not make it.
- Where the two fields hold different standards of evidence, SAY SO
EXPLICITLY rather than reconciling them.
- End with three questions I should ask a collaborator from that field.
Task: Write the memo.
```
**Say why the notes and not the papers:** *"Rung 4 is the only rung where I do not hand it the literature. It gets what I understood. If my understanding is thin, the memo will be thin — and that is the correct outcome, because a memo that is better than my understanding is a memo I cannot defend."*
### The deliberate failure — the 90 seconds that make the point
Ask it to add one sentence: *"Add a line quantifying how much algorithmic bias increases health disparities."*
It will reach for Obermeyer's headline: remedying the disparity would raise the share of Black patients flagged for extra help **from 17.7% to 46.5%** `[59]`.
**Stop and run the Compatibility Check on that number:**
1. **What bar did it clear?** It is a counterfactual simulation at one threshold, in one commercial algorithm, on one health system's patients — not a population estimate.
2. **What is the unit of evidence?** An observational audit of a deployed system. Field A has GRADE, RoB 2 and PRISMA for grading such things `[35]``[37]``[38]`; ask what field B applied.
3. **What is the base rate of it holding up?** Unmeasured here. In neighbouring fields, replication of headline effects runs **36% to 62%** `[31]``[32]``[33]``[34]`.
4. **Is the number comparable?** No — it is specific to that algorithm's proxy target. The paper's own mechanism sentence says so: *"the algorithm predicts health care costs rather than illness."* `[59]`
5. **Did I search where field B publishes?** Not yet.
**Land the session here:** *"The number is real, the paper is excellent, and the sentence would still have been wrong — because it would have travelled into my evidence base without the conditions it was true under. That is what 'different evidence standards' means in practice. It goes in the memo as a **question for my collaborator**, not as a finding."*
Then close the loop: *"And when I publish this work — cite across fields, frame within one. On 128,950 manuscripts including rejections, interdisciplinary **references** raised acceptance; interdisciplinary **topic language** lowered it."* `[8]`
---
## 7. Fallbacks — assume something breaks
| If this fails | Do this |
|---|---|
| Connected Papers is down, or you have used your 5 free graphs | Show yesterday's screenshot. Say the limit out loud — "5 graphs per month on the free tier" `[43]` — because attendees planning a review need to hear it. |
| Inciteful is unreachable, or the domain redirect confuses the room | Show the pre-run connector screenshot and say: "the old address, inciteful.xyz, now redirects to incitefulmed.com — that is a thing that happens to free tools, and it is a reason to record the URL and the date on your map." `[44]` |
| The assistant refuses, stalls or times out on Prompt 1 | Paste your pre-run glossary. The check (§3) is the demonstrable part and needs no live model at all. |
| A DOI actually fails to resolve | The alternate path — §5a scripts it. Slow down, show it, give the 2025 gradient `[21]`, then run 5b anyway: existence was never the whole check. |
| A DOI resolves to a *different* paper than the annotation describes | This is the best possible outcome. Slow down and show it. It is identifier-level failure and annotation-level failure in one — exactly what rung 3's check exists to catch. |
| Both annotations in 5b match their papers | Say so, plainly: *"Today the list survived the check. The 39–77% support rates `[66]` say that will not always happen, and the check costs seven minutes either way."* Do not manufacture a failure. |
| You are running long | Cut §4 step 2a (the similarity graph) and keep the connector. Never cut §5b. |
| Someone asks for a tool recommendation | *"I am not recommending any of these. Slides 17 and 18 give you what each one documents about itself; pick on your corpus, your institution's rules and your budget."* |
---
## 8. Presenter notes
**On the field pair.** If your audience is entirely non-clinical, keep the *shape* and swap the content: two fields that study the same object, have grown up separately, and share at least three words with different meanings. That shape is what makes the demo teach. Do not swap in your own field pair on the day, though — you will defend the AI's glossary instead of testing it, and the room will notice.
**On the tone.** Rung 1 is the one people will try first, and it is the one that will fool them, because a fluent glossary feels like understanding `[23]`. Repeat the check more than feels necessary.
**On what not to do.** Do not run rung 4 before rung 3, or rung 3 before rung 2. The whole session is an argument that each rung's *check* is what makes the next rung's input safe, and the demo has to model that order or it teaches the opposite.
**On honesty about AI.** Somewhere in the Q&A, say the uncomfortable part: across 41.3 million papers (measured through 2025), AI-augmented scientists publish **3.02×** more and are cited **4.84×** more — while AI adoption shrank the collective range of topics studied by **4.63%** and engagement between scientists by **22%** `[16]`. *"The tool that makes it cheap for me to cross into your field is the same tool that is making the field-crossing rarer overall. The counterweight is rung 4's last instruction: go and talk to someone."*
---
## 9. Close (1 minute)
> *"Four rungs. Translate, map, read, synthesise. Each one produces something you keep: a glossary, a landscape, eight papers, a memo. Each one has one check, and you cannot skip a rung, because each rung's check is what makes the next rung's input safe to use. Fill in **row 6** of your canvas before you leave — the check column first. And notice what rung 4 ends with: a question for a human being. Twenty-five minutes of AI got us to a better question. It did not get us an answer, and it was never going to."*
---
*AI for Researchers · Session 6: AI for Multidisciplinary Research · Landscape as of August 2026 · Footnote markers `[n]` resolve in `sources.md`.*