2-Minute Recap · The Whole Series
The Series So Far
- Session 1 — The Capability/Failure Map for Research Tasks: every task sits in Safe, Unreliable or Dangerous. Rule 2: rarity moves tasks rightward.
- Session 1 — The 2026 Tool Landscape Taxonomy: Tier 1 — General chatbots · Tier 2 — Research-specific tools · Tier 3 — Deep-research agents. Today: what the taxonomy left out.
- CRIT — Context / Role and register / Instructions and constraints / Task, then iterate.
- Session 2 — The Literature-Discovery Tool Matrix: six citation-grounded tools, compared from their own documentation.
- Session 3 — The Summary-Trust Triage: when a summary obliges you to reopen the paper.
- Session 4 — The Publisher AI-Policy Comparison Table: what to disclose, and where. Today's medical segment sends you back to it.
- Session 5 — The Personal AI Workflow Canvas: one page, one project — with a deliberately empty row 7. Today you fill it, and the canvas is finished.
- Session 6 — The Cross-Field Orientation Ladder: four rungs into someone else's field. Today you meet what they actually use when you get there.
Today (Session 7): the AI that is not a chatbot — the systems reshaping five disciplines, what each one really does, and how to track yours after this series ends.
Literacy Foundation — The Key Slide
Specialized AI Is Not a Better Chatbot. It Is a Different Object.
- It was already decisive before the chatbots arrived. The 2024 Nobel Prize in Physics went to Hopfield and Hinton "for foundational discoveries and inventions that enable machine learning with artificial neural networks" [2]; the Chemistry prize, the same week, went half to Baker "for computational protein design" and half jointly to Hassabis and Jumper "for protein structure prediction". [1]
- It outputs a domain object, not prose — a coordinate file, a candidate crystal, a proof term, a coded transcript. You cannot judge it by reading it, which is exactly how Sessions 3 and 6 taught you to judge prose.
- It is scored by somebody else. CASP [66], Matbench Discovery [47], an FDA premarket review [7], an IMO paper [49]. Blinded, external, repeated — the thing Tier 1 tools have never had.
- And it usually hands off to a verifier that is not software: a wet lab, a telescope, a proof kernel, a second human coder. That verifier is the whole safety story [3].
So Session 1's Tier 1 / 2 / 3 sorted tools by how general they are. Today's question is different and better: what checks it, and did you run the check?
Series Artefact — The Reusable Part of Today
Five Questions to Ask Any Field-Specific AI Tool
- What exactly does it output, and what does its own paper say it does not do? Every honest landmark system publishes a limitations section. AlphaFold 3's names a 4.4% chirality violation rate and hallucinated order in disordered regions. [19]
- Who scored it, and were they independent of the people who built it? A vendor benchmark and a blinded community assessment are not the same evidence. [66] [47]
- What is the gap between the benchmark and the deployment? Trial-matching scored 87.3% on synthetic patients [16] and precision 0.33 on 157 real ones. [17] Same task.
- What is the external verifier, and can you afford to run it? Ninety-six wells, a telescope night, a proof kernel, a second coder. If you cannot run it, you cannot use the output as a finding.
- What happens when the inputs shift? When ECMWF upgraded its physics model, the fine-tuned ML forecasters lost skill while the physics gained it. [37]
These five travel across all five fields, and they are the spine of the tracker on slide 26. Write them on the inside cover of your notebook; the tool names will change by Christmas.
Medicine — Landmark Systems, Regulatory Context
1,524 Authorised Devices, and Almost None of Them Is What You Picture
1,524entries counted by us in the 2026-06-16 download of FDA's AI-Enabled Medical Device List; decisions spanning 1995 to 30 March 2026 [7]
76.4%of those entries have Radiology as lead review panel — our count, same download. FDA publishes none of these numbers as prose [7]
0identified as large-language-model based — and still 0: FDA's sentence is future tense, it "will explore methods to identify and tag medical devices that incorporate foundation models" [7]
- The exemplar of autonomous diagnosis is narrower than its reputation. IDx-DR (2018, the first FDA-authorised autonomous diagnostic AI) is indicated to "automatically detect more than mild diabetic retinopathy". Its own labelling: it "is not intended to detect concomitant diseases", "does not screen for glaucoma", "does not treat retinopathy". [8] One disease, one camera, one referral decision.
- Read the list as FDA writes it — and read the date on your own count. "The list is not a comprehensive resource of AI-enabled medical devices"; entries were keyword-matched from public summaries, and it "will continue to be updated periodically". [7] Re-checked 2026-09-02: both quotes unchanged, still nothing tagged as a foundation model — the zero holds, but those counts are a snapshot.
Medicine — What Changed for Practitioners
The Number That Travels, and the Number That Does Not
87.3% → 0.33Clinical-trial matching, 2024–25 systems. TrialGPT: criterion-matching accuracy 87.3% on 1,015 pairs, screening time down 42.6% — on 183 synthetic patients [16]. Four deployed tools on 157 real tumour-board patients: mean precision 0.33, recall 0.32, and 38% of patients got no trial at all [17]. No comparable real-patient evaluation of a 2026-generation matcher was found — the gap, not the number, is what transfers
36% → 66%, but 31%Imaging. Across 173 CE-certified radiology AI products, the share with any peer-reviewed evidence rose from 36% to 66% — yet only 31% have evidence at the level of clinical decisions, outcomes or cost, and prospective designs went 19% → 16% [11]
- Where the evidence is strong, it is narrow and prospective. Autonomous diabetic-retinopathy screening: pivotal trial sensitivity 87.2%, specificity 90.7% across 10 US primary-care sites [9]; pooled across 82 studies and 887,244 examinations, sensitivity 0.93 / specificity 0.90 [10].
- And the same meta-analysis names the drift: false positives rise with any-DR screening, low-income settings and ungradable images — hence its call for "post-market audits with standardized gradability metrics". [10] A scoping review of 140 studies found only 7 on implementation and 6 on cost. [12]
Medicine — The Caution
Twenty Hours of AI Training Did Not Protect Them
"Physicians demonstrate substantial automation bias when exposed to erroneous LLM recommendations, even with voluntary consultation and prior AI literacy training."
Randomised trial run in 2025 on a GPT-4-class assistant, 44 physicians who had completed a 20-hour AI-literacy course, 264 cases: composite diagnostic accuracy fell from 84.9% to 73.3%, adjusted −14.0 points (95% CI −18.9 to −9.1, P < 0.0001) [15]
- Three randomised trials, all on the 2023–25 cohort, one uncomfortable pattern. Giving physicians an LLM changed diagnostic reasoning by 2 points (95% CI −4 to 8, P = .60) in a 2024 trial [13] and 6.5 points (95% CI 2.7–10.2) in a 2025 GPT-4 trial [14], where the model alone was statistically indistinguishable from the model-plus-physician (−0.9%, P = 0.8) [14]. No RCT of a 2026-generation assistant has replaced these — which is why the effect to plan for is automation bias, a property of use, not of a model version.
- The tools that work best in medicine are the least conversational. PubMed-wide entity and relation search over ~36M abstracts and 6M full texts, with a relation-extraction F-score of 82.0% — free, no login, and it shows you the sentence it drew the relation from. [18]
- Disclosure is already settled elsewhere in this series. Journal and publisher obligations for AI-assisted manuscripts live in Session 4's Publisher AI-Policy Comparison Table — go back to it before you write up anything built with these tools. Device authorisation [7] governs the product; the policy table governs your paper. They are different regimes and you are subject to both.
Biotech — Landmark System, From Its Own Paper
AlphaFold 3: What It Does, and What Its Authors Say It Does Not
Does
- Predicts the joint structure of complexes containing proteins, nucleic acids, small molecules, ions and modified residues — one model, not protein-only [19]
- Beats classical docking without structural inputs (vs Vina, P = 2.27 × 10⁻¹³) [19]
- Underwrites a database of over 200 million predicted structures, free and bulk-downloadable [22]
- Blinded verdict, CASP16: single-domain fold prediction is "nearly solved" [21]
Does not
- Give dynamics: static outputs "as seen in the PDB"; multiple seeds "do not approximate the solution ensemble" [19]
- Guarantee chemistry: 4.4% chirality violation rate; hallucinated order in disordered regions, without AF2's visual tell [19]
- Rank outputs: "model ranking remains a persistent weakness across most groups" [21]
- Come without strings: 30 jobs/day, closed ligand list, 2021-09-30 cutoff (changeable default); weights gated, non-commercial [20]
Every line on both sides comes from the AF3 paper [19], DeepMind's own docs [20], or the independent CASP16 assessment [21] [23] — never from coverage. Quotas as documented 2026-07-29; the gated, non-commercial weights terms re-read 2026-09-02, unchanged. An AF3 model that looks clean is not evidence that it is.
Biotech — What Changed for Practitioners
Design Moved From Reading Proteins to Writing Them — at a Measured Cost
- Three tools, three different jobs. ProteinMPNN writes a sequence for a given backbone — 52.4% native sequence recovery vs 32.9% for Rosetta [24]. RFdiffusion generates the backbone [25]. ESM3 jointly generates sequence, structure and function (discrete annotation tokens) [28] — its generated esmGFP's "58% identity" is to tagRFP, an engineered protein, not wild type; the closest natural relative is eqFP578 at 53% [28].
- The wet lab still decides, and the numbers are public. RFdiffusion binders: 95 designs tested per target across five targets, 19% overall experimental success — "roughly two orders of magnitude" better than the same group's Rosetta method [25]. BindCraft reports 10–100% across twelve targets, averaging 46.3% [26]. Its "<0.1%" Rosetta figure describes older literature, not a matched comparison [26].
- Enzymes are harder, read the units carefully. RFdiffusion2 scaffolds all 41 benchmark sites against 16 for RFdiffusion — but that is in-silico. Experimentally it found actives "testing fewer than 96 sequences"; the best design — round two of 96, with a general base — reached 53,000 ± 5,000 M⁻¹s⁻¹. [27] [70]
The design software is free; the verifier is your consumables budget — question 4 on slide 6. "41 of 41" is computational; "19%" and "46.3%" are experimental. Never compare across that line.
Biotech — Where the Story Is Contested
AI Drug Discovery: State the Claim, Then State Its Own Authors' Caveat
"In Phase I we find AI-discovered molecules have an 80–90% success rate, substantially higher than historic industry averages… In Phase II the success rate is ~40%, albeit on a limited sample size, comparable to historic industry averages."
Jayatunga, Ayers, Bruens, Jayanth & Meier, Drug Discovery Today 2024 — on 24 molecules, 21 successes. The authors call this "early signs of potential", not proof [29]
- The counterweight, printed inside the most positive clinical paper of 2025. The phase 2a trial of a generative-AI-discovered TNIK inhibitor states in its own main text that "AI-discovered drugs have experienced similar levels of phase 2 trial failure as non-AI-discovered drugs, and none has so far progressed through phase 3 trials". [30] Still true, and we checked: the first registered phase 3 (NCT07687459, 320 patients) was still "not yet recruiting" on 2026-09-02. [30]
- Read that trial the way you would read any other. n = 71 across four arms; FVC change +98.4 mL (95% CI 10.9–185.9) at the top dose against −20.3 mL (95% CI −116.1 to 75.6) for placebo. [30] Promising, small, phase 2a.
- A peer-reviewed challenge to the framing itself: the field optimises target discovery while leaving untouched the human-response measurement that actually kills drugs. [68] Same pattern, one layer down: Geneformer and scGPT, zero-shot across five tissue datasets, underperformed highly variable gene selection, scVI and Harmony — and some evaluation data overlapped pretraining, so that is the best case. [31] A 2025 result on 2023–24 models: re-run it before you cite it.
Engineering — Landmark System, and the Published Correction
Simulation Surrogates: Real, Operational, and Systematically Oversold
25 Feb 2025ECMWF's ML forecast system AIFS Single goes operational; the ensemble follows 1 July 2025; both to v2 on 12 May 2026 [36]
5–15%the honest accuracy gain: medium-range error reduction vs the physics-based IFS. The order-of-magnitude win is in cost, not skill [36]
79%of ML-for-fluid-PDE papers claiming to beat a numerical method compared against a weak baseline (60 of 76) [32]
- Read these two together, not side by side. The claim of "speedups of four to five orders of magnitude" comes from a review written by the developers of the method reviewing their own method [34], and rests on inference-versus-solver timings — precisely the comparison class that a Nature Machine Intelligence analysis found weakly baselined in 79% of cases, with reporting and publication bias "widespread" [32]. Treat any order-of-magnitude speedup as unverified until you see the baseline.
- And the failure is not a capacity problem. Physics-informed neural networks "can learn good models for relatively trivial problems" but "easily fail to learn relevant physical phenomena for even slightly more complex problems" — an optimisation failure, not an architecture one. [33]
Engineering — What Changed, and the Caution
"Generative Design" Is Not a Text Prompt — and Your Model Ages
- What commercial generative design takes as input, per the vendor: preserve geometry ("a body you want to include in the final shape") and obstacle geometry ("a body you want to exclude"), plus loads, constraints, materials, manufacturing methods [39]. Verified July 2026: no text prompt in the loop — the vendor blocked our 2026-09-02 re-check and has announced prompt-driven CAD: re-date this one.
- Text-to-CAD is a separate, younger thing, and at v1.4.4 (Aug 2026) its vendor still says so: "still experimental", designs that "aren't manufacturable or safe". [40] A first solid, not an assembly.
- Code assistants: two dated baselines. On "Early-2025 AI", 16 experienced developers on their own repos were 19% slower (95% CI +2% to +39%) while believing they were 20% faster [42]; on 2022-era Copilot, across 89 security scenarios, ~40% of 1,689 programs were vulnerable [41]. Historical measurements on named generations — quote them with the year, or not at all.
- The caution that generalises to every field. When ECMWF upgraded its physics model the fine-tuned ML forecasters — GraphCast, Aurora, AIFS v1.1 — lost skill while the physics gained it; ECMWF stopped running them in real time. [37]
[42] is still a preprint (carried from Session 5); neither code number has been re-measured on the 2026 models. Dataset shift is not hypothetical: it is why a surrogate validated this year needs revalidating when its upstream data changes — and why open benchmarks exist. [38]
Physical Sciences — Where the Story Is Contested
Materials Discovery: the Claim, and the Peer-Reviewed Rebuttal
2.2M · 380kThe claim (Nature, 2023). GNoME predicted 2.2 million structures, ~380,000 stable — "an order-of-magnitude expansion in stable materials known to humanity" [43]. A companion lab reported synthesising novel materials over 17 days [45]
0 · 0The rebuttals (peer-reviewed, 2024). "Scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility… we have yet to find any strikingly novel compounds" [44]. On the autonomous lab, re-examining all 43 products: "no new materials have been discovered in that work" [46]
- Both critics also say the method itself is sound. Predictions order cations a real solid would disorder — 0 K DFT has no entropy — and "automated Rietveld analysis… is not yet reliable" [44] [46]; yet the same paper calls the approach "sound" [44].
- The correction landed in the journal. An Author Correction (Nature, 2026) restates the result as "36 compounds from a set of 57 targets", drops four inconclusive compounds and one in the training data, retitles "novel materials" to "inorganic materials", and concedes they were "new to the prediction platform, not necessarily new to science". [69]
The live answer is a leaderboard, not a headline: Matbench Discovery's peer-reviewed hedge is that models "effectively and cheaply pre-screen" candidates; the crisper "triaging steps" line is the project's GitHub README, not the journal. [47] Snapshot, 2026-09-02: 42 models, led by EquiformerV3+DeNS-OAM (F1 0.931), added April 2026 — stale the moment it is read.
Physical Sciences — What Changed for Practitioners
In Observational Physics, AI Is Not the Discovery. It Is the Only Way to Get to It.
"Rubin issued 800,000 alerts the night of 24 February… a system expected to eventually produce up to seven million alerts per night… scientists rely on a network of intelligent software platforms known as brokers. These systems use machine learning algorithms to filter, sort, and classify the alerts."
NSF–DOE Vera C. Rubin Observatory, first scientific alerts, 25 February 2026 — seven full-stream brokers plus two down-stream services, as its data-products page describes them; alerts world-public [48]
- Note what the ML is doing, and what it is not. It is triage and classification at a rate no human pipeline can approach — seven million events a night, a public alert within about two minutes of the shutter. It does not decide what is real, and it does not do the physics. Every interesting alert still becomes a follow-up observation.
- That is the pattern across the physical sciences, and it is the healthy one: the model produces candidates, the instrument produces evidence. Where the two get confused — as in the materials dispute on the previous slide — the correction takes years. [44] [46]
- The transferable habit: ask what your field's alert stream is, then ask who is allowed to re-examine the classifier's output. Rubin's answer is "anyone" — the alerts are world-public. [48] Very few AI claims in any field survive that test.
Mathematics — Landmark System, Stated Exactly
Theorem Provers: the Only Field Whose Verifier Cannot Be Fooled
- What AlphaProof did at IMO 2024. It proved three of the five non-geometry problems in Lean — from statements manually formalised by humans — each needing 2–3 days of test-time RL. Geometry went to another system; both combinatorics problems went unsolved. Combined: 28/42. [49]
- And what it did not do. It did not read the problems as posed or work in contest time, and its authors write that the bespoke training "is likely beyond the reach of most academic research groups". [49]
- The working tool is the proof assistant, not the prover. Lean's Mathlib holds 286,514 theorems from 772 contributors [50] (2026-09-02; it advances daily), and a machine-checked proof is the one AI output in this deck that you do not have to trust — the kernel accepts it or it does not.
- Which sets up the sharpest contrast in the session. In July 2025 two labs reported gold-medal IMO scores from undisclosed systems. The IMO's statement that week: "the IMO cannot validate the methods, including the amount of compute used or whether there was any human involvement, or whether the results can be reproduced." [52] The 2026 olympiad (Shanghai, July 2026) has been held since; that is still the IMO's last word.
The IMO records the AI companies as a fringe event testing "closed-source AI models… privately" [52]. Benchmarks are not neutral ground either: in June 2026 one research-mathematics benchmark shipped a v2 "addressing errors in 42% of problems". [51]
Social Sciences — Landmark Result, and the Correction
Annotation at Scale Works. Plugging the Labels Into a Regression Does Not.
"Direct use of surrogate labels in downstream statistical analyses leads to substantial bias and invalid confidence intervals, even with high surrogate accuracy of 80–90%."
Egami, Hinck, Stewart & Wei, NeurIPS 2023 — because model errors are non-random and correlate with the covariates you are regressing on [54]
- The result that started it — read the comparator and the date. Across four datasets and 6,183 texts, ChatGPT in 2023 beat MTurk crowd workers by ~25 points at under $0.003 per annotation [53]. But absolute accuracy was 59–83%, and the trained-annotator result is on intercoder agreement, not accuracy. This is not "LLMs beat expert coders".
- The result that qualifies it. Coding 34 constructs with GPT-4: κ ≥ 0.70 for 25 of them — but only with the best prompting strategy chosen per construct, and "no single method consistently outperforms the others". Zero-shot ranged κ = 0.91 to κ = 0.11. [56] A 2025 paper on a 2023 model.
- The finding most likely to outlive the model. GPT-4 struggled most on exactly the constructs human coders struggled to agree on — human–human κ itself ran 0.24–0.87. [56] Manifest codes transfer; interpretive ones do not. A better model raises κ; it cannot make an ambiguous construct unambiguous.
- Reproducibility is a separate problem. "An experiment carried out with a given model often yielded different results when repeated a few weeks later", and deprecation makes reproducibility "nearly impossible". [55] Record model version and date, every time.
Social Sciences — Where the Story Is Contested
"Silicon Samples": a Debate Fought on 2022–24 Models — and Still Unsettled
The claim
- GPT-3, 2023. Conditioned on sociodemographic backstories, a model can "accurately emulate response distributions from a wide variety of human subgroups" — the authors' term: algorithmic fidelity [57]
- Preprint, revised June 2026. Self-report grounding works: across 1,052 Americans, interview-only, survey-only and combined agents reached 83%, 82% and 86% of participants' own two-week test–retest consistency, vs 74% demographics-only [60]
The rebuttals
- GPT-3.5, 2024. Means match; the rest does not. "Less variation in responses than in the real surveys, and regression coefficients often differ significantly"; "the same prompt yields significantly different results over a 3-month period" [58]
- Independent replication, 2025. Models "cannot replace research subjects", showing "strong bias and a low variance on each topic" — and "this bias randomly varies from one topic to the next" [59]
Read this pair as a live argument, not a verdict: every number was measured on a 2022–24 model and nobody has re-run either side on the 2026 generation. What survives is structural: matching the mean is not matching the distribution, social science lives in the variance, and a topic-random bias [59] cannot be corrected without the human data you were avoiding. [60] is still an unpublished preprint (v3, checked 2026-09-02) whose revision shows a structured survey (82%) buys what a two-hour interview (83%) does.
Humanities — The Best-Verified Result in This Deck
A Scroll Nobody Could Open, Read End to End
"PHerc. 1667, sealed since the eruption of Vesuvius in 79 AD, has been virtually unwrapped and read from beginning to end… roughly 1.4 metres of papyrus and around twenty-two columns of Greek."
Vesuvius Challenge, 25 June 2026 — X-ray tomography, geometric reconstruction, ML ink detection; every reading transcribed and reviewed by papyrologists; data openly licensed, code on GitHub, preprint posted [61]
- It is exemplary because of how it is reported. The project states its own limit in the announcement: "Because the papyrus is damaged, the readings are fragmentary, with gaps where the surface is lost." [61] Compare that to the two disputes earlier in this deck.
- The everyday humanities tool is handwritten text recognition, and its vendor publishes indicative error bands rather than a headline: character error rate under 2% is publication-ready, 2–5% good, 5–15% for medieval manuscripts — needing 15–30 pages of hand-corrected ground truth for a single hand, and 50–100 for a multi-hand collection. [62] Vendor-stated, not independently evaluated.
- The caution for text-based humanities. A review synthesising five existing benchmarks of historical competence — not new experiments — reports "a collapse in LLM performance as assessments move from contaminated to decontaminated datasets, from Western to global knowledge domains, and from multiple-choice questions to open-ended responses". [63] Recognition is not reasoning.
Demo Preview — You Choose Two, Live
Five Field Demos Are Loaded. We Run the Two You Vote For.
- A · Medicine. Entity-and-relation search across the whole of PubMed, then the same question put to a current chatbot — which may well be right, and still cannot show you where it got it. Free, no login. [18]
- B · Biotech. A public protein sequence into a structure-prediction server: read the confidence colouring, find the disordered region, and see why "it looks clean" is not evidence. [19] [20]
- C · Engineering. Text-to-CAD on a part a first-year could draw, then on a part an engineer would actually need — and the vendor's own words for what happens next. [40]
- D · Physical sciences & mathematics. A false theorem, proved fluently by a chatbot and pasted into Lean in the browser: the kernel rejects it however good the prose is. Then a true one, and the kernel accepts. The only demo in this series with a verdict. [50]
- E · Social sciences & humanities. Ten synthetic survey responses, a four-code book, LLM coding against human coding, Cohen's κ computed live — and what a high κ still does not license you to do. [54] [56]
Each runs about five minutes. All five are written out in full in the demo script with prompts, expected outcomes and fallbacks, so the three we do not run are still yours. No real patient, participant or unpublished material appears in any of them — the sample data is synthetic and printed in the script.
Closing the Series
Seven Sessions, Seven Artefacts, One Habit
"The proliferation of AI tools in science risks introducing a phase of scientific enquiry in which we produce more but understand less."
Messeri & Crockett, Nature 2024 — on illusions of understanding, and the scientific monocultures they hide [4]
- Every artefact in this series was a defence against that one sentence. The Capability/Failure Map and CRIT (S1), the Literature-Discovery Tool Matrix (S2), the Summary-Trust Triage (S3), the Publisher AI-Policy Comparison Table (S4), the Personal AI Workflow Canvas and the Four Reproducibility Guardrails (S5), the Cross-Field Orientation Ladder (S6) — and today, the Field-Frontier Tracker. Seven pages. None of them names a product.
- The habit under all seven is a single question: what would show me this is wrong, and did I go and look? That is the whole series, and it is not new — it is just what research already was.
- Two things to do this week. Fill row 7 of your canvas for one real project. Then build a Field-Frontier Tracker for your own field, and put a date on it.
- And one thing to notice. AI's benefits in science are not evenly distributed — fields with higher proportions of women or Black scientists have so far reaped fewer of them. [5] Whoever in your field builds the tracker and shares it decides who gets to keep up. Make it a shared document.