AI for Researchers · Visual Deck
Five fields, five landmark systems, and the one question that tells you whether any of them is safe to use
{{Presenter Name}}
Landscape as of August 2026
2-Minute Recap · The Whole Series
flowchart LR s1(["1 · Map · Taxonomy
· CRIT"]):::rest --> s2(["2 · Tool
Matrix"]):::rest s2 --> s3(["3 · Summary-
Trust Triage"]):::rest s3 --> s4(["4 · Policy
Table"]):::rest s4 --> s5(["5 · Canvas —
row 7 empty"]):::rest s5 --> s6(["6 · Orientation
Ladder"]):::rest s6 --> s7(["7 · Field-Frontier
Tracker
today"]):::today classDef today fill:#ab7d22,stroke:#d9b36c,color:#14132b,font-size:19px classDef rest fill:#23204c,stroke:#8f8cb8,color:#b9b7d6,font-size:19px
Today (Session 7): the AI that is not a chatbot — the systems reshaping five disciplines, what each one really does, and how to track yours after this series ends.
Row 7 of Session 5's canvas gets filled today, and the canvas is finished.
Learning Objectives
Full verbatim wording in the reference deck and curriculum
Section 01
Six sessions on tools that write. This one is about tools that predict a structure, propose a compound, or refuse a proof — and why they fail in a completely different direction.
Literacy Foundation — The Key Slide
flowchart LR M["a domain model
not a chat box"]:::model --> O["a domain object, not prose
a coordinate file · a candidate crystal ·
a proof term · a coded transcript"]:::obj --> V["a verifier that is not software
a wet lab · a telescope ·
a proof kernel · a second human coder"]:::ver classDef model fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef obj fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef ver fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px
Tier 1/2/3 sorted tools by how general they are. Today's question is different and better: what checks it, and did you run the check?
Series Artefact — The Reusable Part of Today
The spine of the Field-Frontier Tracker (slide 26). Write them on the inside cover of your notebook; the tool names will change by Christmas.
Section 02
The only field on today's tour where a regulator already decides what the software is allowed to claim — and where the published numbers fall furthest between the benchmark and the clinic.
Medicine — Landmark Systems, Regulatory Context
The autonomous exemplar is narrower than its reputation: IDx-DR — one disease, one camera, one referral decision; "not intended to detect concomitant diseases" [8]. And FDA's own words: the list is "not a comprehensive resource" — those counts are a snapshot [7].
Medicine — What Changed for Practitioners
No comparable real-patient evaluation of a 2026-generation matcher was found — the gap, not the number, is what transfers. The meta-analysis itself calls for "post-market audits with standardized gradability metrics" [10] [12].
Medicine — The Caution
No RCT of a 2026-generation assistant has replaced these — plan for automation bias, a property of use, not of a model version. The least conversational tools work best: PubMed-wide entity search, F-score 82.0%, shows you the sentence [18]. Disclosure for your paper: Session 4's Policy Table.
Section 03
The field with the clearest win of the decade, the most disciplined published limitations, and the most contested commercial promise.
Biotech — Landmark System, From Its Own Paper
Every line from the AF3 paper [19], DeepMind's own docs [20], or the independent CASP16 assessment [21] [23] — never coverage. An AF3 model that looks clean is not evidence that it is.
Biotech — What Changed for Practitioners
The design software is free; the verifier is your consumables budget — question 4. "41 of 41" is computational; "19%" and "46.3%" are experimental. Never compare across that line. ESM3 generates sequence, structure and function jointly [28].
Biotech — Where the Story Is Contested
"In Phase I… an 80–90% success rate… In Phase II… ~40%, albeit on a limited sample size, comparable to historic industry averages."
Drug Discovery Today 2024 — on 24 molecules, 21 successes; the authors call this "early signs of potential", not proof [29]
"AI-discovered drugs have experienced similar levels of phase 2 trial failure… and none has so far progressed through phase 3 trials."
Printed inside the most positive clinical paper of 2025 (phase 2a, n = 71: FVC +98.4 mL top dose vs −20.3 mL placebo). First registered phase 3, NCT07687459, still "not yet recruiting" on 2026-09-02 [30]
A peer-reviewed challenge to the framing itself [68] — and one layer down: Geneformer and scGPT, zero-shot, underperformed highly variable gene selection, scVI and Harmony [31]. A 2025 result on 2023–24 models: re-run it before you cite it.
Section 04
One operational triumph, one measured literature-wide bias, and a word — "generative design" — that does not mean what the audience assumes it means.
Engineering — Landmark System, and the Published Correction
The "four to five orders of magnitude" speedup review was written by the method's own developers [34]; PINNs "easily fail… for even slightly more complex problems" — an optimisation failure [33]. Treat any order-of-magnitude speedup as unverified until you see the baseline [32].
Engineering — What Changed, and the Caution
Dataset shift is not hypothetical: when ECMWF upgraded its physics, the fine-tuned forecasters lost skill and ECMWF stopped running them [37] — why open benchmarks exist [38]. [42] is still a preprint.
Section 05
Where two disciplines built the best answer anyone has to AI hype: they let outsiders re-examine the claim. One of them got a correction. The other could not even try.
Physical Sciences — Where the Story Is Contested
The live answer is a leaderboard, not a headline: Matbench Discovery — models "effectively and cheaply pre-screen" candidates [47]. Snapshot, 2026-09-02: 42 models, led by EquiformerV3+DeNS-OAM (F1 0.931), added April 2026 — stale the moment it is read.
Physical Sciences — What Changed for Practitioners
flowchart LR SKY["Rubin Observatory
800,000 alerts on night one —
up to seven million per night"]:::io --> BR["brokers — seven full-stream,
two down-stream: ML filter,
sort and classify"]:::model --> CAND["candidates
public alert ~two minutes
from the shutter"]:::hot --> FU["follow-up observation
= the evidence"]:::ver classDef io fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef model fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef hot fill:#23204c,stroke:#d9b36c,color:#eceaf8,font-size:19px classDef ver fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px
The model produces candidates; the instrument produces evidence — where the two get confused, the correction takes years [44] [46].
Rubin's alerts are world-public [48] — ask who is allowed to re-examine your field's classifier. Very few AI claims survive that test.
Mathematics — Landmark System, Stated Exactly
flowchart LR CB["a chatbot's fluent proof
true or false — you cannot tell by reading it"]:::n --> L["formalised
in Lean"]:::n --> K{"proof
kernel"}:::kern K -- accepts --> OK["machine-checked — the one AI output
you do not have to trust"]:::ok K -- rejects --> NO["rejected —
however good the prose"]:::bad classDef n fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef kern fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef ok fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px classDef bad fill:#1c1a3f,stroke:#c14b58,color:#ef9a9a,font-size:19px
The contrast: July 2025's gold-medal claims came from undisclosed systems — the IMO "cannot validate the methods… or whether the results can be reproduced" [52], still its last word. And one research-maths benchmark shipped a v2 "addressing errors in 42% of problems" [51].
Section 06
The field where the tool is a general model doing a specialised job — and where accuracy on the label is the easy half of the problem.
Social Sciences — Landmark Result, and the Correction
ChatGPT in 2023 beat MTurk crowd workers by ~25 points at under $0.003 — but absolute accuracy was 59–83% [53]. Manifest codes transfer; interpretive ones do not [56] — and repeated runs drift, so record model version and date, every time [55].
Social Sciences — Where the Story Is Contested
A live argument, not a verdict: every number is from a 2022–24 model, and nobody has re-run either side on the 2026 generation. Matching the mean is not matching the distribution — and [60] is still an unpublished preprint (v3, checked 2026-09-02): a structured survey (82%) buys what a two-hour interview (83%) does.
Humanities — The Best-Verified Result in This Deck
"PHerc. 1667, sealed since the eruption of Vesuvius in 79 AD, has been virtually unwrapped and read from beginning to end… roughly 1.4 metres of papyrus and around twenty-two columns of Greek."
Vesuvius Challenge, 25 June 2026 — X-ray tomography, geometric reconstruction, ML ink detection; every reading reviewed by papyrologists; data openly licensed, code on GitHub; the announcement names its own limit: readings "fragmentary, with gaps where the surface is lost" [61]
The caution: "a collapse in LLM performance" from contaminated to decontaminated datasets, Western to global domains, multiple-choice to open-ended [63]. Recognition is not reasoning.
Series Artefact · Today's Canonical Workflow
The verifier and the venue are the whole artefact — a field's frontier is not a list of tools. Rebuild it for your own field; it stays true when every product name changes. Full five-column table in the reference deck.
Series Artefact — Session 5's Canvas, Row 7
Three rules: if you cannot name the verifier, the tool is not ready for your project · a benchmark number is not a deployment number — 87.3% became 0.33 on real patients [16] [17] · re-date the row every six months; inputs drift [37].
Demo Preview — You Choose Two, Live
~Five minutes each; all five written out in the demo script with prompts, expected outcomes and fallbacks — the three we do not run are still yours. All sample data is synthetic.
Closing the Series
"The proliferation of AI tools in science risks introducing a phase of scientific enquiry in which we produce more but understand less."
Messeri & Crockett, Nature 2024 — on illusions of understanding, and the scientific monocultures they hide [4]
This week: fill row 7 of your canvas for one real project, then build a Field-Frontier Tracker for your own field — and put a date on it.
References · Cross-cutting · Medicine · Biotech
References · Biotech · Engineering · Physical Sciences
References · Physical Sciences & Mathematics · Social Sciences
References · Social Sciences & Humanities · Frontier Venues