← series index

# Demo Script — Session 7: Discipline Deep-Dives: Specialized AI by Field

*AI for Researchers · presenter script · Landscape as of August 2026*

**Runtime:** ~25 minutes of the 60-minute session (following ~20 minutes of literacy foundation, leaving ~15 for Q&A).
**Format:** **audience poll.** Five field demos are prepared in full below. You run the **top two**.
**Artefacts it exercises:** the Five Questions (slide 6) and the Field-Frontier Tracker (slide 26).
**Canvas link:** every demo fills **row 7** of Session 5's Personal AI Workflow Canvas — and with it, the canvas is finished.

**Budget:** 3 min poll and framing · 2 × ~10 min (a 5-minute walkthrough, then ~5 minutes on what it does *not* do and questions) · 2 min close.

---

## 0. What this demo block is, and what it deliberately is not

This is **not** a tour of five tools. It is the same question asked five times, in five fields, of software that is not a chatbot:

> *"What does this thing actually output, who checks it, and did I run the check?"*

The three demos you do not run are still the audience's — every one of them is written out below with verbatim prompts, expected outcomes and a fallback, and the handout points here.

**Say this out loud before the poll:**

> *"Two of these five will be new to almost everyone in the room, including me. That is the point of the session. I am not going to demonstrate expertise in five fields — I am going to demonstrate the same three-minute interrogation in five fields, and you are going to be able to run it on Monday in yours."*

**A standing rule, stated once, applied to all five:** no real patient, participant, unpublished sequence, or unpublished manuscript goes into any of these tools, live or otherwise. Every input below is either **synthetic and printed in this script**, or **already public** with a stable primary URL.

---

## 1. Preparation — the day before, not the morning of

| # | Do this | Why |
|---|---|---|
| 1 | **Run all five demos end to end and save every output** (screenshots + exported files). These become the fallbacks in §8. | Four of the five depend on a live third-party service. Assume one is down. |
| 2 | Re-check each tool still exists and still behaves as scripted. All five were verified live on **2026-07-29**; the fast-moving endpoints (Mathlib counts, Zoo Design Studio version and output format, FDA list wording, the phase-3 registration) were re-checked **2026-09-02**. | The engineering tool moved *during* the writing of this script, then moved again — version, agent name and output format — six weeks later. See §5. This is not a hypothetical risk. |
| 3 | Submit the **AlphaFold Server** jobs (§4) the day before. | Jobs are queued, not instant. Never wait on a queue on camera. |
| 4 | Copy every prompt below into a scratch file. **Do not retype them live.** | Every prompt here is load-bearing; a typo changes what you are demonstrating. |
| 5 | Open the poll (§2) on one screen and a browser with **five pre-loaded tabs** on the other. | You have 3 minutes between the vote and the first click. |
| 6 | Pre-compute the κ values for Demo E (§7) from your own run, on the **named model you will use live**, and write them and the model name on a card. | If the live output differs — and it will — you need the pre-run numbers to compare against. That difference *is* a finding, in either direction. |
| 7 | Have `sources.md` open. | Someone will ask "where does 19% come from?" It is `[25]`. |

---

## 2. The poll (~3 minutes)

**Mechanics.** Any show-of-hands or webinar-poll widget. Ask the question by *field*, not by tool name — attendees who do not recognise a tool name will not vote for it, which biases you towards the famous ones.

> **"Which two of these five do you most want to see? Vote for two."**
> **A** — Medicine & clinical research
> **B** — Biotechnology & life sciences
> **C** — Engineering
> **D** — Physical sciences & mathematics
> **E** — Social sciences & humanities

**If the vote is close or split five ways,** take A and D. They are the two most different from each other: A is a specialised search engine over real literature, D is the only demo in the entire series whose output carries a verdict that no amount of fluency can move. Say why you chose them.

**If one field takes >60% of the vote,** run it first and take the runner-up second — do not run two demos from the same field even if the room asks.

---

## 3. Demo A — Medicine & clinical research: PubTator 3.0 (~5 minutes)

**Tool:** PubTator 3.0, NCBI/NLM. `https://www.ncbi.nlm.nih.gov/research/pubtator3/`
**Access:** free, **no login**, no institutional subscription. Verified live 2026-07-29.
**Why this tool:** it is the cleanest example in medicine of a specialised system that is *not* conversational. It does entity recognition and relation extraction over ~36M PubMed abstracts and 6M PMC full texts, with a published relation-extraction F-score of **82.0%** — the 2024 paper's own benchmark figure, not a claim about your query `[18]`.

### Step 1 — Frame it (30 s)

> *"This is not a chatbot and it will not write you a paragraph. It has read all of PubMed and tagged six kinds of thing — genes and proteins, chemicals, diseases, species, genetic variants, cell lines — and twelve kinds of relation between them. Watch what that buys you that fluency does not."* `[18]`

### Step 2 — Search an entity (1 min)

Type into the PubTator 3.0 search box, verbatim:

```
metformin
```

**Expected outcome:** a ranked result list; matched entities highlighted in titles and snippets; a facet/entity panel offering related entities. Click any result to open the abstract with **coloured entity annotations in place**.

**Say:** *"Every one of those colours is a normalised identifier, not a string match. 'Metformin', 'Glucophage' and its CID all collapse to the same entity — which is exactly what keyword search cannot do."*

### Step 3 — The move that matters: a relation query (1.5 min)

In the search box, verbatim:

```
metformin AND "colorectal cancer"
```

Then open one result and point at the highlighted **chemical** and **disease** entities in the same sentence.

**Expected outcome:** papers where both normalised entities co-occur, with the sentence carrying the relation visible on screen.

**Say the key line:** *"I can see the sentence the claim came from. That is the whole difference. A chatbot gives me a claim; this gives me a claim **and its provenance in one screen**, so checking it costs me two seconds rather than twenty minutes."*

### Step 4 — The contrast, and the "does not" (1.5 min)

Open a current general chatbot in the next tab — ChatGPT, Claude, Gemini or Kimi; **name the one you are using and the date**, because the answer depends on both — and paste, verbatim:

```
List the five most-cited papers on metformin and colorectal cancer risk.
Give the full citation and PMID for each.
```

**Expected outcome:** a fluent, confident, well-formatted list — and on the current generation, with search switched on, **the PMIDs will very often resolve.** That is the expected path, not a disappointment. **Do not predict that it is wrong.** Pick **one** PMID, paste it into PubMed live, and let the room watch what happens.

**Run the check both ways, out loud:**

- **It resolves and the paper says what was claimed.** Good. Now ask the question that matters: *"which sentence, in which paper, did that come from?"* Neither you nor the model can point at it without going back to the text.
- **It resolves but the paper does not support the claim.** This is the common 2026 failure and the reason link-resolution is no longer the test.
- **It does not resolve at all.** Rarer than it was; still worth naming when it happens.

**Say the key line:** *"Notice that I could not tell those three cases apart from the screen. The chatbot may be completely right — and it is still the case that **the only way to know is to leave the chatbot**. PubTator had already shown me the sentence. That is not a difference in accuracy; it is a difference in **provenance**, and provenance is what survives a model upgrade."*

Then state PubTator's limits, out loud, so the room does not over-trust the other side either:

- It finds **relations that were stated in text**. It is not evidence of a causal relationship, and it does not appraise study quality.
- The 82.0% F-score is a **benchmark** number `[18]`. Your query is not the benchmark.
- Coverage is PubMed abstracts plus the PMC **open-access** subset — not all full text.

### Step 5 — Land it on the artefact (30 s)

> *"Five Questions, slide 6. Output: annotated entities and relations. Scored by: a published F-score on a public corpus. Verifier: **the sentence itself, which is on my screen**. That is why this one is cheap to check and a chatbot answer is not."*

**Fallback A:** if the site is slow or down, use the documented REST endpoint in a browser tab — it returns JSON directly and was verified live on 2026-07-29:
`https://www.ncbi.nlm.nih.gov/research/pubtator3-api/search/?text=metformin`
If the network is gone entirely, use your saved screenshots and run **Step 4 only**, which needs no specialised tool.

---

## 4. Demo B — Biotechnology & life sciences: AlphaFold Server (~5 minutes)

**Tool:** AlphaFold Server. `https://alphafoldserver.com/` (Google account required; **30 jobs/day**; **non-commercial** terms) `[20]`
**Companion:** AlphaFold Protein Structure Database, `https://alphafold.ebi.ac.uk/` — free, no login, >200 million predictions `[22]`. Verified live 2026-07-29.

### The two inputs — both public, neither unpublished

Copy the FASTA from UniProt's REST endpoint rather than retyping (a single-residue typo silently changes the demo):

| Job | Protein | Copy from | Expect |
|---|---|---|---|
| 1 | **Human lysozyme C** (148 aa) | `https://rest.uniprot.org/uniprotkb/P61626.fasta` | A tight, well-packed fold, **high confidence nearly everywhere** |
| 2 | **Human α-synuclein** (140 aa) | `https://rest.uniprot.org/uniprotkb/P37840.fasta` | An intrinsically disordered protein: **low confidence, and shapes that look like structure but are not** |

> **Never paste an unpublished or commercially sensitive sequence into a public prediction server.** Say this on camera. The AlphaFold Server terms are non-commercial and the weights are gated at Google DeepMind's discretion `[20]` — that is a licensing fact your institution's contracts office may care about more than you do.

### Step 1 — Show the pre-run pair (2 min)

Open your two **pre-run** results side by side. Switch the colouring to **pLDDT / model confidence**.

**Say:** *"Same software, same length, same five minutes of compute. One of these is a triumph and one of these is a warning, and the only thing telling them apart is the colour."*

- Lysozyme: blue/high confidence throughout. Cross-check against the experimental structure in the AlphaFold DB entry for P61626.
- α-Synuclein: orange/yellow, low confidence, long ribbons. **This protein is genuinely disordered in solution.** The model still draws *something*.

### Step 2 — The "does not", read from the paper (1.5 min)

Put the AlphaFold 3 paper's own limitations on screen and read them:

- Outputs are **static** structures "as seen in the PDB"; multiple seeds "do not approximate the solution ensemble" `[19]`.
- A measured **4.4% chirality violation rate** `[19]`.
- Spurious order in disordered regions — and, unlike AlphaFold 2, **without the visual tell** that used to make it obvious `[19]`.
- Independently, CASP16: single-domain fold prediction is "nearly solved", but "**model ranking remains a persistent weakness across most groups**" `[21]`.

**Say the key line:** *"Confidence colouring is the verifier you get for free. It is also the only one you get for free — everything else on this screen is a hypothesis that costs bench time."*

### Step 3 — Submit one live, then leave it (1 min)

Submit job 1 live so the room sees the interface and the job limit. **Do not wait for it.** Move on.

### Step 4 — Land it on the artefact (30 s)

> *"Output: coordinates plus a per-residue confidence. Scored by: CASP, blinded, about a hundred groups, and none of them work for the people who built it `[66]`. Verifier: the wet lab — published de novo binder success is **19%** in one landmark paper `[25]` and **10–100%** across twelve targets in another `[26]`. The software is free. The verifier is your consumables budget."*

**Fallback B:** the entire demo works from the **AlphaFold DB** (`alphafold.ebi.ac.uk`) with no login and no queue — search `P61626` and `P37840` and show the same two confidence pictures. If the network is gone, your saved screenshots carry Steps 1, 2 and 4 unchanged; only the live submission is lost.

---

## 5. Demo C — Engineering: Zoo Design Studio, Text-to-CAD (~5 minutes)

**Tool:** Zoo Design Studio, browser build — `https://app.zoo.dev/`
**Access note, and read it before the day:** the standalone `text-to-cad.zoo.dev` page forwards into the browser build `[40]`. Design Studio was at **v1.3.7** (page updated 2026-07-22) when this script was written; on re-check **2026-09-02 it was v1.4.4, page updated 2026-08-28**, the conversational agent is now branded **Zookeeper**, and the documented output has shifted: CAD-generation results "can include editable **KCL** project files", while "producing formats such as **STEP** requires a separate modeling and export step" `[40]`. **Check the version and the output format again the week you present.**

> **Use this as the opening line.** *"This tool moved while I was writing the script for it — and it moved again between two drafts of this script, twice in six weeks, including what file it hands you at the end. That is the single most useful thing I can teach you about tracking a field's frontier: the URL in your notes is stale before the slide deck is."*

### Step 1 — Establish what "generative design" already meant (1 min)

Before touching the AI tool, state the definition, in the incumbent vendor's own words — inputs are "**hold out areas, preserved areas, loads, and constraints**", plus materials and manufacturing methods `[39]`.

**Say:** *"Commercial generative design has existed for years and it has **no text prompt anywhere in the loop**. It is multi-objective shape and topology search under stated physics. What we are about to use is a different, much younger thing wearing a similar name."*

### Step 2 — The easy prompt (1 min)

Paste verbatim:

```
A 20 mm cube with a 10 mm diameter hole through the centre of one face,
passing all the way through.
```

**Expected outcome:** a correct, clean solid, in seconds. Rotate it. Then point at **what you are actually handed** — as of 2026-09-02 that is an editable **KCL** project, with **STEP** export a separate step `[40]`. Do the export live if it is one click; if it is not, say so.

**Say:** *"That is real. It is a B-rep solid, not a mesh or a picture — it goes into a CAD package and it is manufacturable. Note that what lands in my hands is a **project file**, not a STEP file, and that changed between two versions of this tool. The handoff format is part of the tool, and it drifts."*

### Step 3 — The honest prompt (2 min)

Paste verbatim:

```
A motor mount bracket for a NEMA 17 stepper motor: 5 mm thick aluminium,
four M3 clearance holes on the standard 31 mm bolt circle, a 22 mm
central bore, and a 90-degree flange with two M4 slots for adjustment.
```

**Do not predict the outcome. Run the check.** These builds change monthly, and the honest instruction is the one you want the room to copy: **measure a feature on screen** — bolt circle 31 mm, bore 22 mm, flange 90°, slots that are slots and not holes.

**Then narrate whichever you got:**

- **Something is wrong** (a bolt circle that is not 31 mm, slots rendered as holes, a flange at the wrong angle, a bore off concentricity). The commonest result, and the easiest to teach from.
- **Every dimension you measure is right.** Say so — *"and here is why that does not settle it"* — then keep measuring: tolerance, wall thickness at the fillet, whether the slots clear an M4 head, whether it is manufacturable in the stated 5 mm aluminium. **A part that passes four checks has passed four checks.**

**Say the key line:** *"I cannot tell by looking. Neither can you — and notice that this stayed true even in the run where it was right. In this field the verifier is a dimension check and then a physical fit, and the failure mode is not a hallucinated citation, it is a **hole in the wrong place** — which is a failure mode nothing earlier in this series prepared you for, because you cannot read it off the screen."*

Then quote the vendor against itself — this is fair, because it is what their own documentation says, still, on 2026-09-02:

- The model is "**still experimental**" and "some results may not be as good as you expect" `[40]`.
- Their conversational agent, **Zookeeper**, "can make mistakes - it may misunderstand intent, produce incorrect geometry, or suggest designs that **aren't manufacturable or safe**" `[40]`.

**Worth saying out loud:** the product page has moved to confident marketing language ("production-ready CAD"), while the FAQ still carries both warnings. *"When a vendor's front page and its own FAQ disagree about how much to trust the tool, believe the FAQ."*

### Step 4 — Widen to the field, and land it (1 min)

> *"Two numbers to take away, and **both come with a year attached — say the year, every time you quote them.** On what the authors call 'Early-2025 AI', sixteen experienced open-source developers on their own repositories were measured **19% slower** while believing they were 20% faster `[42]` — a preprint, but a careful one. And on **2022-era GitHub Copilot**, across 89 security-relevant scenarios, about **40%** of generated programs were vulnerable `[41]`. Neither has been re-run on the 2026 models, so neither is a claim about the tool you used this morning. What has not changed is the structure: engineering's verifier has always been validation and verification, and AI does not reduce that burden — it increases the number of candidate designs arriving at it."*

**Fallback C:** the browser build needs an account and a network. If either fails, show your saved screenshots of both prompts — the pedagogical content is entirely in the **comparison**, not in the live generation. As a zero-dependency substitute, run Step 1 and Step 3 as a *thought* experiment against the incumbent-vendor definition `[39]` and go straight to Step 4.

---

## 6. Demo D — Physical sciences & mathematics: Lean 4 + LeanSearch (~5 minutes)

**Tools:** Lean 4 Web, `https://live.lean-lang.org/` (browser, **no install, no login**) and LeanSearch, `https://leansearch.net/` (free semantic search over Mathlib) `[50]`. Both verified live 2026-07-29.

**Why this is the best demo in the series:** it is the only one where the verdict comes from something that cannot be argued with. **Note the inversion from earlier versions of this script:** the honest prior for the current model generation is that a chatbot's proof of an easy true theorem **compiles**. So the demonstration is no longer "watch the model fail" — it is "watch the kernel decide", and the way to show that unmistakably is to hand it something **false**.

> **Read the scope limit before you start.** The maintainers scope the web editor to "smallish" snippets and state that "serious Lean code development and larger projects are considered out-of-scope" `[50]`. We are well inside that.

### Step 1 — Set up (30 s)

Open `https://live.lean-lang.org/` and put this on the first line:

```lean
import Mathlib
```

Wait for the orange progress bar to clear. **Say:** *"That is Mathlib loading — 286,514 theorems from 772 contributors, every one machine-checked `[50]`. Nothing in it is there because someone was persuasive."*

### Step 2 — PRIMARY PATH: ask a current chatbot to prove something that is false (1 min)

In a chatbot tab (ChatGPT, Claude, Gemini or Kimi — say which one you are using and note the date), paste verbatim:

```
Write a Lean 4 proof, using Mathlib, of the theorem that for every
natural number n, n * n is even. Output only the Lean code.
```

`n * n` is even only when `n` is. **The statement is false, and no proof of it exists.**

**Two things can happen, and both are the demo — say which one you got:**

- **It writes the proof anyway.** Most likely something confident and well-shaped. Paste it into Lean (Step 3).
- **It refuses, and tells you the statement is false.** Say so cheerfully — *"good, it caught it; now watch me not have to take its word for that"* — and paste the printed attempt below instead. **You have not lost the demo; you have gained a second one.**

**The printed attempt** — plausible Lean, fluent, and unprovable because the theorem is false:

```lean
theorem even_sq (n : ℕ) : Even (n * n) := by
  induction n with
  | zero => simp
  | succ k ih =>
    rw [Nat.succ_mul, Nat.mul_succ]
    exact ih.add (Even.add_one (by simpa using ih))
```

A one-line version, if you want the error to arrive instantly:

```lean
theorem even_sq (n : ℕ) : Even (n * n) := by decide
```

### Step 3 — Paste it into Lean and let the kernel answer (1.5 min)

Paste under `import Mathlib`. **Say nothing.** Let the red squiggles arrive.

**Expected outcome — this one is not a gamble:** errors. The exact message varies with the model and the Mathlib version — an unsolved goal, a failed tactic, a `decide` that cannot decide an unbounded statement — but **a false theorem cannot be proved, so there is no version of the current model generation that makes this green.**

**Say the key line:** *"That is the entire session in one screen. It was fluent, it was confident, it was formatted correctly — and the kernel does not care. **I did not have to be a mathematician, and I did not have to out-argue the model.** I had a verifier, and I ran it."*

### Step 4 — The control: a true theorem, and a second kind of tool (1.5 min)

Now show the other half, so the room does not leave thinking the tools are useless.

Ask the same chatbot, verbatim:

```
Write a Lean 4 proof, using Mathlib, of the theorem that for every
natural number n, n * (n + 1) is even. Output only the Lean code.
```

**Expected outcome, current generation: it compiles.** Paste it in and let it go green. Say so plainly — *"and that is the honest state of play in 2026; on an easy true statement the model is usually right."* **If it does not compile** — it sometimes will not — you have the older, harder version of this demo, so narrate the error and move on; the lesson does not change.

Then show the specialised alternative. Go to `https://leansearch.net/` and search, verbatim:

```
n * (n + 1) is even
```

**Expected outcome:** the relevant Mathlib lemma near the top, with its exact current name. Or let Mathlib find it itself:

```lean
theorem even_mul_succ (n : ℕ) : Even (n * (n + 1)) := by
  exact?
```

`exact?` reports the lemma it used. Accept the suggestion. **Squiggles clear.**

**Say the closing line — this is the one to get right:** *"Two AI tools, one job each: one wrote a candidate proof, one searched a library of verified ones. Notice what actually settled both of them. **Not which tool was cleverer — the kernel.** The model was right on the true theorem and would have been believed on the false one. The verifier is what made the difference between those two cases visible to me in one second, and your field has a verifier too. It is probably an experiment, and it probably costs money."*

### Step 5 — Land it, with the frontier caution (30 s)

> *"AlphaProof proved three of five non-geometry IMO problems in 2024 — from statements **manually formalised by human experts**, each needing **2–3 days** of test-time compute, with both combinatorics problems unsolved `[49]`. Then in July 2025, two labs reported gold-medal scores from systems nobody outside them has seen. The IMO's own statement: it '**cannot validate the methods**, including the amount of compute used or whether there was any human involvement, or whether the results can be reproduced' `[52]`. Same week, same competition, completely different epistemic status — and as of the 2025 cycle that statement is still the IMO's own last word on the question. The 2026 olympiad has been held since, in Shanghai; if you present this after a fresh IMO, **check for a new statement before you quote the old one.**"*

**Fallback D:** both services are free and lightweight, but if Lean 4 Web is unavailable, this demo runs perfectly from **screenshots of the red error state (the false theorem) and the green cleared state (the true one)** — the content is the contrast between them, and a screenshot of a red squiggle is still a red squiggle. Keep a copy of the false-theorem attempt and the working true proof in your scratch file.

---

## 7. Demo E — Social sciences & humanities: LLM coding against a codebook (~5 minutes)

**Tools:** any general LLM, plus a spreadsheet. Optionally **Taguette** (`https://www.taguette.org/`, free, open source, BSD, **no AI features**) as the manual-coding contrast `[64]`. Verified live 2026-07-29.
**Data: ten SYNTHETIC responses, written for this script.** They are not from any real study, participant or dataset. Say so on camera.

### The synthetic corpus

Study prompt (fictional): *"What has made it hard to adopt data-management practices in your lab?"*

| # | Synthetic response |
|---|---|
| R1 | "Honestly there just isn't time. I'm on a nine-month contract and every hour I spend tidying files is an hour I'm not writing." |
| R2 | "Nobody ever showed me how. I asked our IT people and they sent me a 40-page PDF." |
| R3 | "We have a shared drive that fills up every March. Then someone deletes things and we never find out what." |
| R4 | "I do it, but I'm not convinced anyone will ever look at any of it. It feels like paperwork for a funder rather than for science." |
| R5 | "The training was fine but it was all about a repository we don't use." |
| R6 | "Time. Always time." |
| R7 | "My PI thinks it's important, so we do it. I'd rather be at the bench, but I can see the argument." |
| R8 | "There's no budget line for storage after the grant ends, so we quietly keep everything on a lab laptop." |
| R9 | "I learned it from a postdoc who left. Since then nobody has kept the wiki up to date." |
| R10 | "It's not that it's hard. It's that it's boring, and boring things slip." |

### The codebook — three manifest codes and one latent code

| Code | Definition |
|---|---|
| **TIME** | Respondent names time, workload or competing priorities as a barrier. |
| **TRAIN** | Respondent names absent, unclear or mismatched training or knowledge transfer. |
| **INFRA** | Respondent names infrastructure, storage, tooling, budget or institutional policy. |
| **AMBIV** | *Latent.* Respondent expresses ambivalence about whether the practice is **worth doing** — compliance without conviction. Not the same as naming a barrier. |

### The human gold standard (two coders, reconciled)

| | TIME | TRAIN | INFRA | AMBIV |
|---|---|---|---|---|
| R1 | ✔ | | | |
| R2 | | ✔ | | |
| R3 | | | ✔ | |
| R4 | | | | ✔ |
| R5 | | ✔ | ✔ | |
| R6 | ✔ | | | |
| R7 | ✔ | | | ✔ |
| R8 | | | ✔ | |
| R9 | | ✔ | ✔ | |
| R10 | | | | ✔ |

### Step 1 — Code it (1.5 min)

Paste verbatim, in one message, with the corpus and codebook appended:

```
Context: I am coding open-ended survey responses for a qualitative study.
Role and register: Act as a second coder applying a fixed codebook. You are
not designing the codebook and you may not add, merge or rename codes.
Instructions and constraints:
- Apply ONLY these four codes: TIME, TRAIN, INFRA, AMBIV, using the
  definitions given verbatim below.
- A response may take zero, one or several codes.
- Output a table: response number, then one column per code, marked 1 or 0.
- After the table, and separately, list any response you found genuinely
  ambiguous and say in one sentence why.
- Do not explain your reasoning inside the table and do not quote the
  responses back to me.
Task: Code all ten responses.

[PASTE THE CODEBOOK]
[PASTE R1-R10]
```

**Expected outcome:** a clean 10 × 4 table, produced in seconds.

### Step 2 — Score it against the humans (1.5 min)

Put the model's table next to the gold table in a spreadsheet and compute **Cohen's κ per code** over the 10 decisions for that code:

`κ = (p_o − p_e) / (1 − p_e)`, where `p_o` is observed agreement and `p_e` is chance agreement from the marginals.

**Do not promise a result. Compute the κ and then read whichever of these two you got — the demo works either way, and the second one is now at least as likely as the first.**

- **Path 1 — AMBIV falls apart.** High agreement on **TIME**, **TRAIN** and **INFRA** (typically κ ≈ 0.7–1.0) and markedly worse on **AMBIV**. The classic failure is the model coding R1 or R2 as AMBIV — reading a *barrier* as *doubt about value* — and missing it in R10, where the ambivalence is carried by the word "boring". **Say:** *"the interpretive code is where it breaks, and it breaks exactly where your two human coders would have argued."*
- **Path 2 — it handles AMBIV too.** On the current model generation this happens regularly: κ near 1.0 across all four codes. **Do not treat this as the demo failing — it is the demo arriving at the more important half.** Say: *"Good. It got the hard one. Now watch me show you why that is not permission to do the thing you were about to do"* — and go straight to **Finding 2** in Step 3, which is the real payload of this demo and does not depend on the κ at all.

> **Use your pre-run numbers from §1.6 here.** If the live run disagrees with your rehearsal, show both. *"Same prompt, same ten responses, different day, different numbers"* is one of the most valuable things that can happen in this demo — and it is a published finding, not an accident: "an experiment carried out with a given model often yielded **different results when repeated a few weeks later**" `[55]`. **A run-to-run difference is a finding whichever direction it goes.**

### Step 3 — The two findings that matter (1.5 min)

**Finding 1 — where it worked and where it did not, with a citation and a vintage.** In a peer-reviewed evaluation of **GPT-4** across 34 constructs, κ ≥ 0.70 was reached for 25 of them — **but only when the best prompting strategy was chosen per construct**, and "no single method consistently outperforms the others"; zero-shot ranged from **κ = 0.91 to κ = 0.11** `[56]`. And GPT-4 struggled most with exactly the constructs human coders struggled to agree on — human–human κ itself running **0.24–0.87** `[56]`.

**Say:** *"That is a 2025 paper measuring a 2023 model, and the current models score better — I would expect them to. What I would not expect to change is the **shape**: AMBIV is not hard because the model is stupid, it is hard because it is interpretive, and your two human coders would have fought about it too. A better model raises the κ. It does not make an ambiguous construct unambiguous."*

**Finding 2 — the one that will cost someone a paper, and the reason this demo works even when the model wins.** Do **not** skip this, ever — and if the κ came back high, **this is your demo**, not a coda to it:

> *"Suppose κ came back at 0.85 — or, on today's models, 0.95 — and you are delighted. If you now put those labels into a regression as if they were data, you get '**substantial bias and invalid confidence intervals, even with high surrogate accuracy of 80–90%**' `[54]`. Read that clause again: **the bias is a consequence of high accuracy, not a consolation for low accuracy.** Model errors are not random — they correlate with exactly the covariates you are regressing on, so a better model gives you a more confidently wrong coefficient. The fix is design-based: you hand-code a **random** subset with a sampling probability **you** fixed in advance, and you use an estimator built for surrogate labels. That decision has to be made before you annotate, not after — and **no model release will retire it.**"*

### Step 4 — Land it, with the field's contested edge (30 s)

> *"And if anyone tells you that you can skip collecting human data entirely: the evidence they will cite, on both sides, was measured on **2022–24 models** — GPT-3 for the original claim `[57]`, GPT-3.5 for the first rebuttal `[58]`. Nobody has re-run it on this year's models, so treat it as the record of an argument rather than a verdict. What the rebuttals found was **structural**: synthetic respondents match the **mean** and lose the **variance** — 'less variation in responses than in the real surveys, and regression coefficients often differ significantly' `[58]`, replicated independently as 'strong bias and a low variance on each topic', with the bias '**randomly varying from one topic to the next**' `[59]`. Social science lives in the variance, and a topic-random bias is one you cannot correct without the human data you were trying not to collect."*

**Fallback E:** this demo has the fewest dependencies of the five — the corpus, codebook and gold standard are printed above, so it can be run entirely on a spreadsheet from your pre-run model output. If you lose the model as well, hand the room the corpus, ask *them* to code R4, R7 and R10, and let the disagreement in the room make the point about latent codes for you. That version needs no software at all and is often better.

---

## 8. Fallbacks at a glance

| Demo | First fallback | Zero-network fallback |
|---|---|---|
| **A** Medicine | REST endpoint `pubtator3-api/search/?text=metformin` | Saved screenshots + run the chatbot-contrast step only |
| **B** Biotech | AlphaFold DB (`alphafold.ebi.ac.uk`, no login, no queue) | Saved confidence screenshots for P61626 and P37840 |
| **C** Engineering | Saved screenshots of both prompts | Incumbent-vendor definition `[39]` + the two numbers in Step 4 |
| **D** Maths | Screenshots of the false theorem's error state and the true theorem's cleared state | Same — the contrast between the two is the content |
| **E** Social | Pre-run model table + spreadsheet | Ask the room to code R4, R7 and R10 themselves; Finding 2 stands alone |

**Universal fallback:** if two services fail at once, stop demoing and run the **Five Questions** (slide 6) out loud against whichever field the room voted for, using the Field-Frontier Tracker (slide 26). That is the transferable content; the tools were only ever the illustration.

---

## 9. Closing the demo block (~2 minutes)

**Say:**

> *"Five fields, five tools, and the same three questions each time: what does it output, who scored it, and what is the verifier. Notice which demos were comfortable. The two where I could check the answer in seconds — the sentence on the PubTator screen and the Lean kernel — are the two where the AI was least able to mislead me, **and notice that this had nothing to do with how good the models were today.** In several of these the model was right. The comfort came from the **verifier**, not from the model, and the verifier is the part that does not get a new version number every six weeks."*

> *"So the last thing to write down is not a tool name. It is row 7 of your canvas: the one bounded object you will hand to a specialised tool on your next project, and the name of the thing that will tell you it is wrong. If you cannot name the second one, the tool goes on your reading list, not into your project."*

Then hand over to Q&A, and point at the handout: **all five demos are written out there**, including the three the room did not vote for.

---

*AI for Researchers · Session 7: Discipline Deep-Dives: Specialized AI by Field · Landscape as of August 2026*