← series index
# Sources — Session 7: Discipline Deep-Dives: Specialized AI by Field
*AI for Researchers · dated, annotated reference list · compiled 2026-07-29 · modernization pass 2026-08-21 · fast-moving endpoints re-checked 2026-09-02 · Landscape as of August 2026*
Entry format (one line per source, numbering matches the `[n]` footnote markers used in `slides.html`, `demo-script.md` and `handout.md`):
`- [n] Author/Org (Year). *Title*. URL — accessed YYYY-MM-DD. [type: peer-reviewed | policy | primary-doc | news] — one-line note on what claim(s) it grounds.`
Allowed `type` values:
- `peer-reviewed` — peer-reviewed paper or arXiv/preprint literature
- `policy` — official publisher/funder/institutional/regulator policy page
- `primary-doc` — primary tool/vendor/observatory documentation (not a blog or listicle)
- `news` — recent news/benchmark article; informs tool discovery only, never the sole ground for a factual claim
**Evidence rule for this session.** This deck tours five fields no single presenter is expert in, so the rule is stricter than usual and is stated on the slides themselves:
1. **Every named landmark system is described from its own primary paper or its own documentation** — never from press coverage, and never from memory.
2. **Every landmark system gets a "does not" as well as a "does",** taken from the same source. Where the authors state a limitation, that wording is quoted.
3. **Where a field's AI story is contested, both sides are cited** — GNoME/A-Lab (`[43]`–`[46]`), AI drug discovery (`[29]`–`[30]`), ML surrogates (`[32]`–`[34]`), silicon sampling (`[57]`–`[59]`). The contest is the content; the press-release version is not on these slides.
4. **Regulatory claims cite the regulator**, dated (`[7]`, `[8]`).
5. **Preprints are labelled as preprints** on the slide that uses them (`[42]`, `[60]`).
6. Vendor blogs and listicles informed tool discovery; none is load-bearing. Vendor documentation is cited only for facts about that vendor's own product (`[20]`, `[39]`, `[40]`, `[62]`, `[64]`).
**Tally (recounted programmatically from the `[type:]` tags, updated 2026-08-08 with the addition of [70]; the 2026-08-21 modernization pass added and removed none, so the counts stand):** **70 sources — 50 peer-reviewed/preprint, 18 primary-doc, 3 policy/regulator** (one entry, [22], is tagged `primary-doc + peer-reviewed` and appears in both counts; 50 + 18 + 3 − 1 = 70). No source in this deck is `news`: for a session whose whole risk is misdescribing a landmark system, nothing short of a primary paper, a regulator's own file or a maintainer's own documentation was allowed to carry a claim. Field tags in bold at the head of each section.
---
## References
### A. Cross-cutting — what "specialized AI" is, and what it has already done
- [1] The Royal Swedish Academy of Sciences / NobelPrize.org (2024). *The Nobel Prize in Chemistry 2024*. https://www.nobelprize.org/prizes/chemistry/2024/summary/ — accessed 2026-07-29. [type: primary-doc] — Exact citation wording and prize shares: one half to David Baker "for computational protein design", the other half jointly to Demis Hassabis and John Jumper "for protein structure prediction" (1/2, 1/4, 1/4).
- [2] The Royal Swedish Academy of Sciences / NobelPrize.org (2024). *The Nobel Prize in Physics 2024*. https://www.nobelprize.org/prizes/physics/2024/summary/ — accessed 2026-07-29. [type: primary-doc] — Awarded jointly to John J. Hopfield and Geoffrey Hinton "for foundational discoveries and inventions that enable machine learning with artificial neural networks".
- [3] Wang, H., Fu, T., Du, Y., et al. (2023). *Scientific discovery in the age of artificial intelligence*. Nature 620, 47–60. https://doi.org/10.1038/s41586-023-06221-2 — accessed 2026-07-29. [type: peer-reviewed] — The standing cross-disciplinary review of AI in the research pipeline; grounds the framing that specialised AI is a set of task-specific methods, not a chatbot.
- [4] Messeri, L., & Crockett, M. J. (2024). *Artificial intelligence and illusions of understanding in scientific research*. Nature 627, 49–58. https://doi.org/10.1038/s41586-024-07146-0 — accessed 2026-07-29. [type: peer-reviewed] — "produce more but understand less"; scientific monocultures; the closing caution of the series.
- [5] Gao, J., & Wang, D. (2024). *Quantifying the use and potential benefits of artificial intelligence in scientific research*. Nature Human Behaviour 8, 2281–2292. https://doi.org/10.1038/s41562-024-02020-5 — accessed 2026-07-29. [type: peer-reviewed] — AI use and benefits widespread and growing especially since 2015; a gap between AI education and application; disciplines with higher proportions of women or Black scientists reap fewer benefits.
- [6] The Royal Society (2024). *Science in the age of AI*. https://royalsociety.org/news-resources/projects/science-in-the-age-of-ai/ — accessed 2026-07-29. [type: policy] — First-party report from a scholarly organisation, drawing on research activities with more than 100 scientists; four recommendations on infrastructure, usability, open science and oversight capacity.
### B. **Medicine & clinical research**
- [7] U.S. Food and Drug Administration (2026). *Artificial Intelligence-Enabled Medical Devices* (the AI-Enabled Medical Device List and its downloadable file). https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices — accessed 2026-09-02 (counts from the 2026-07-29 fetch of the **2026-06-16 download**). [type: policy] — **All three counts on slide 8 are from that one download and are explicitly framed as a snapshot on the slide.** Page content current as of 2026-06-16: 1,524 listed entries, final decisions from 1995-09-29 to 2026-03-30; the radiology share (1,164 of 1,524 = 76.4%) is computed from FDA's own downloadable file, not published by FDA as a percentage. Grounds FDA's verbatim caveat: "The list is not a comprehensive resource of AI-enabled medical devices"; and FDA's statement that it "will explore methods to identify and tag medical devices that incorporate foundation models encompassing a wide range of AI systems, from large language models (LLMs) to multimodal architectures". **Re-check 2026-09-02:** the caveat, the "this list will continue to be updated periodically" sentence and the foundation-model sentence are all unchanged, and the foundation-model sentence is still **future tense** — FDA has not yet tagged any device as foundation-model or LLM-based, so the "0 identified as LLM-based" headline stands. A programmatic re-count was **not** possible: the live page renders its table progressively and no total or "content current" date survived our fetch, so the 1,524 and 76.4% figures are deliberately left as the dated 2026-06-16 figures rather than silently refreshed. **If a future download shows any entry tagged as incorporating a foundation model, slide 8's headline and its rhetorical point both change and must be rewritten.**
- [8] U.S. Food and Drug Administration (2018). *De Novo Classification Request Decision Summary, DEN180001 (IDx-DR)*. https://www.accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf — accessed 2026-07-29. [type: policy] — Exactly what the first FDA-authorised autonomous diagnostic AI is indicated to do, and its labelled limits: "IDx-DR is only designed to detect diabetic retinopathy… is not intended to detect concomitant diseases", "does not screen for glaucoma", "does not treat retinopathy". Pivotal n=900; sensitivity 87.4%, specificity 89.5%.
- [9] Abràmoff, M. D., Lavin, P. T., Birch, M., Shah, N., & Folk, J. C. (2018). *Pivotal trial of an autonomous AI-based diagnostic system for detection of diabetic retinopathy in primary care offices*. npj Digital Medicine 1, 39. https://doi.org/10.1038/s41746-018-0040-6 — accessed 2026-07-29. [type: peer-reviewed] — Prospective, 10 US primary-care sites, 900 enrolled: sensitivity 87.2% (95% CI 81.8–91.2), specificity 90.7% (95% CI 88.3–92.7), imageability 96.1%.
- [10] Wang, Z., et al. (2025). *Systematic review and meta-analysis of regulator-approved deep learning systems for fundus diabetic retinopathy detection*. npj Digital Medicine, 19 Dec 2025. https://doi.org/10.1038/s41746-025-02223-8 — accessed 2026-07-29. [type: peer-reviewed] — 82 studies, 887,244 examinations, 25 devices, 28 countries: pooled sensitivity 0.93 / specificity 0.90 per patient; false-positive rate rises with any-DR screening, low-income settings and ungradable images; authors call for "post-market audits with standardized gradability metrics".
- [11] Antonissen, N., Tryfonos, O., Houben, I. B., Jacobs, C., de Rooij, M., & van Leeuwen, K. G. (2026). *Artificial intelligence in radiology: 173 commercially available products and their scientific evidence*. European Radiology. https://doi.org/10.1007/s00330-025-11830-8 — accessed 2026-07-29. [type: peer-reviewed] — 173 CE-certified radiology AI products from 90 vendors; products with any peer-reviewed evidence rose 36% → 66% across 639 papers, but only 31% have any clinical-decision/outcome/cost evidence, and prospective designs did not improve (19% → 16%).
- [12] Lawrence, R., et al. (2025). *Artificial intelligence for diagnostics in radiology practice: a rapid systematic scoping review*. eClinicalMedicine 83, 103228. https://pubmed.ncbi.nlm.nih.gov/40474995/ — accessed 2026-07-29. [type: peer-reviewed] — 8,013 articles screened, 140 included, only 7 on implementation and 6 on cost; "there is a paucity of evidence in real-world settings, supporting cautiousness in how AI is perceived (e.g., as a complementary tool, not a solution)".
- [13] Goh, E., Gallo, R., Hom, J., et al. (2024). *Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial*. JAMA Network Open 7(10), e2440969. https://doi.org/10.1001/jamanetworkopen.2024.40969 — accessed 2026-07-29. [type: peer-reviewed] — **2024 trial, GPT-4-era assistant.** 50 physicians, 6 vignettes: LLM arm 76% vs 74% control, adjusted difference 2 percentage points (95% CI −4 to 8; P=.60); the LLM alone scored 16 points higher than the control group (95% CI 2–30; P=.03).
- [14] Goh, E., et al. (2025). *GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial*. Nature Medicine, 5 Feb 2025. https://doi.org/10.1038/s41591-024-03456-y — accessed 2026-07-29. [type: peer-reviewed] — **2025 trial, GPT-4 (named in the title — quoted verbatim, never renamed).** 92 physicians, five expert-developed vignettes: LLM arm +6.5 percentage points (95% CI 2.7–10.2, P<0.001), but no significant difference between LLM-augmented physicians and the LLM alone (−0.9%, 95% CI −9.0 to 7.2, P=0.8); authors state the result "should be validated in real clinical practice".
- [15] Qazi, I. A., Ali, A., Khawaja, A. U., Akhtar, M. J., Sheikh, A. Z., & Alizai, M. H. (2026). *Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial*. NEJM AI 3(5), published 2026-04-23. https://doi.org/10.1056/AIoa2501001 — accessed 2026-07-29. [type: peer-reviewed] — 44 physicians who had completed a 20-hour AI-literacy course, 264 cases: the arm shown deliberately erroneous LLM suggestions scored 73.3% against 84.9% for the arm shown error-free suggestions — adjusted −14.0 percentage points (95% CI −18.9 to −9.1, P<0.0001). Both arms were offered LLM output; the contrast is erroneous vs error-free, not LLM vs no LLM. **Model generation:** the trial ran in 2025 on a GPT-4-class assistant. Secondary reporting names it as ChatGPT-4o, but the NEJM AI full text returned HTTP 403 to our fetcher on 2026-09-02 and the medRxiv preprint likewise, so the deck states the **vintage** rather than asserting the build. None of the three physician RCTs ([13], [14], [15]) has been repeated on a 2026-generation assistant.
- [16] Jin, Q., Wang, Z., Floudas, C. S., et al. (2024). *Matching patients to clinical trials with large language models* (TrialGPT). Nature Communications 15, 9074. https://doi.org/10.1038/s41467-024-53081-z — accessed 2026-07-29. [type: peer-reviewed] — Retrieval recalls >90% of relevant trials using <6% of the collection; criterion-level matching accuracy 87.3% on 1,015 manually reviewed pairs; screening time reduced 42.6%. Evaluated on three cohorts of 183 **synthetic** patients — it ranks and explains trials, it does not determine eligibility. **2024 system**; paired on the slide with [17] as a 2024–25 benchmark-vs-deployment baseline. We found no comparable real-patient evaluation of a 2026-generation matcher, and the slide says so rather than implying the 0.33 is a current measurement.
- [17] Kempf, E., et al. (2025). *A prospective pragmatic evaluation of automatic trial matching tools in a molecular tumor board*. npj Precision Oncology, 27 Jan 2025. https://doi.org/10.1038/s41698-025-00806-y — accessed 2026-07-29. [type: peer-reviewed] — 157 consecutive real patients, four publicly available matching tools: mean precision 0.33, recall 0.32; 2.19 trials proposed per patient on average and 38% of patients got none. "We recommend that experts supervise the results."
- [18] Wei, C.-H., Allot, A., Lai, P.-T., et al. (2024). *PubTator 3.0: an AI-powered literature resource for unlocking biomedical knowledge*. Nucleic Acids Research 52(W1), W540–W546. https://doi.org/10.1093/nar/gkae235 — accessed 2026-07-29. [type: peer-reviewed] — ~36M PubMed abstracts + 6M PMC open-access full texts, updated weekly; >1.6 billion entity annotations across six entity types; 33 million relations across 12 relation types; BioREx relation-extraction F-score 82.0%. Free, no login. Grounds the medicine demo.
### C. **Biotechnology & life sciences**
- [19] Abramson, J., Adler, J., Dunger, J., et al. (2024). *Accurate structure prediction of biomolecular interactions with AlphaFold 3*. Nature 630, 493–500. https://doi.org/10.1038/s41586-024-07487-w — accessed 2026-07-29. [type: peer-reviewed] — What AF3 predicts (joint structures of complexes containing proteins, nucleic acids, small molecules, ions and modified residues) and, from the paper's own "Model limitations" section, what it does not: static structures "as seen in the PDB", multiple seeds do not approximate the solution ensemble, a 4.4% chirality violation rate, and spurious order ("hallucinations") in disordered regions that lack the visual tell AF2 gave. Outperforms classical docking (vs Vina, P = 2.27 × 10⁻¹³); antibody–antigen gains require up to 1,000 seeds versus a typical 5. (A second P value against RoseTTAFold All-Atom appeared in an earlier draft of this file; it could not be relocated in the paper and has been removed rather than carried forward unverified.)
- [20] Google DeepMind (2026). *AlphaFold Server — Frequently Asked Questions*; *google-deepmind/alphafold3* repository and *Model Parameters Terms of Use*. https://alphafoldserver.com/faq (accessed 2026-07-29) · https://github.com/google-deepmind/alphafold3 · https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md — weights terms re-read and unchanged 2026-09-02. [type: primary-doc] — **Re-check note:** alphafoldserver.com/faq is a JavaScript application that returns no readable text to a fetcher, so the **30 jobs/day quota, the 5,000-token job size, the closed ligand list and the 2021-09-30 template default are carried forward with their 2026-07-29 date and are labelled as such on slide 12.** The weights terms were re-read directly on 2026-09-02 and are verbatim unchanged: parameters "only available for non-commercial use", granted "at its sole discretion", and "must not" be used "in connection with any commercial activities, including research on behalf of commercial organizations". Original annotation follows. — Current access model: Server allows 30 jobs/day, a 5,000-token maximum job size, a fixed closed ligand list, and a PDB template cutoff of 2021-09-30 that is a **default** users can change, disable, or override by uploading their own templates; the local code accepts arbitrary SMILES ligands and covalent bonds. Code is Apache-2.0; **weights are gated and non-commercial**, granted "at Google DeepMind's sole discretion", and may not be used "in connection with any commercial activities, including research on behalf of commercial organizations".
- [21] Yuan, R., Zhang, J., Kryshtafovych, A., Schaeffer, R. D., Zhou, J., Cong, Q., & Grishin, N. V. (2025). *CASP16 Protein Monomer Structure Prediction Assessment*. Proteins: Structure, Function, and Bioinformatics. https://doi.org/10.1002/prot.70031 — accessed 2026-07-29. [type: peer-reviewed] — The independent community assessment: "the problem of single-domain protein fold prediction is nearly solved — no target folds were incorrectly predicted across all Evaluation Units"; remaining challenges are truncated sequences, irregular secondary structures and interaction-induced conformational changes; "model ranking remains a persistent weakness across most groups"; CASP15 → CASP16 progress "subtle".
- [22] EMBL-EBI & Google DeepMind (2026). *AlphaFold Protein Structure Database*; with Varadi, M., et al. (2024). *AlphaFold Protein Structure Database in 2024*. Nucleic Acids Research 52, D368–D375. https://alphafold.ebi.ac.uk/ · https://doi.org/10.1093/nar/gkad1011 — accessed 2026-07-29. [type: primary-doc + peer-reviewed] — "over 200 million protein structure predictions", freely available, with downloads for the human proteome, 47 other key organisms and Swiss-Prot; the peer-reviewed 2024 figure is "over 214 million".
- [23] Peng, C., Ni, W., Liu, Q., Hu, G., & Zheng, W. (2025). *A comprehensive benchmarking of AlphaFold3 for predicting biomacromolecules and their interactions*. Briefings in Bioinformatics 26(6), bbaf616. https://doi.org/10.1093/bib/bbaf616 — accessed 2026-07-29. [type: peer-reviewed] — Independent third-party benchmark of AF3: grounds the claim that AF3's advantage is uneven across biomolecule classes rather than uniform.
- [24] Dauparas, J., Anishchenko, I., Bennett, N., et al. (2022). *Robust deep learning–based protein sequence design using ProteinMPNN*. Science 378, 49–56. https://doi.org/10.1126/science.add2187 — accessed 2026-07-29. [type: peer-reviewed] — Sequence recovery on native backbones 52.4% versus 32.9% for Rosetta. Designs sequences **for a given backbone**; does not generate backbones and does not predict function.
- [25] Watson, J. L., Juergens, D., Bennett, N. R., et al. (2023). *De novo design of protein structure and function with RFdiffusion*. Nature 620, 1089–1100. https://doi.org/10.1038/s41586-023-06415-8 — accessed 2026-07-29. [type: peer-reviewed] — 95 designs tested per target across 5 targets; **overall experimental binder success rate 19%**, "an increase of roughly two orders of magnitude over our previous Rosetta-based method on the same targets"; best affinities ~30 nM without optimisation.
- [26] Pacesa, M., et al. (2025). *One-shot design of functional protein binders with BindCraft*. Nature 646(8084), 483–492. https://doi.org/10.1038/s41586-025-09429-6 — accessed 2026-07-29. [type: peer-reviewed] — Verbatim: "success rates from 10% to 100%, with an **average success rate of 46.3%**" across 12 targets. **Framing correction:** the paper's "less than 0.1%" figure for Rosetta appears in its introduction as a characterisation of the *prior literature* (citing its refs 2 and 4) — it is **not** a matched head-to-head on the same 12 targets, and must not be presented as one.
- [27] Lauko, A., Ahern, W., et al. (2026). *Atom-level enzyme active site scaffolding using RFdiffusion2*. Nature Methods 23, 96–105 (published online 2025-12-03). https://doi.org/10.1038/s41592-025-02975-x — accessed 2026-07-29. [type: peer-reviewed] — Verbatim: "generates scaffolds for all 41 active sites in a diverse benchmark, compared to 16 using previous methods" — this is an **in-silico** scaffolding result, and "previous methods" means RFdiffusion specifically. Experimentally: active candidates identified "after experimentally testing fewer than 96 sequences in each case", for three catalytic mechanisms. **Catalytic-efficiency figures are reported in an accompanying paper, not measured here — see [70], whose best result (53,000 ± 5,000 M⁻¹s⁻¹) is the k_cat/K_M number carried onto slide 13.**
- [28] Hayes, T., Rao, R., Akin, H., et al. (2025). *Simulating 500 million years of evolution with a language model*. Science 387, eads0018. https://doi.org/10.1126/science.ads0018 — accessed 2026-07-29. [type: peer-reviewed] — ESM3 trained on 3.15B sequences, 236M structures, 539M annotations; generated esmGFP at 58% sequence identity to **tagRFP, an engineered protein**; the closest *wild-type* relative is eqFP578 at 53%. ESM3's "function" track is **discrete annotation tokens** (InterPro / keyword labels), not a predicted activity — so "joint sequence, structure and function generation" must be qualified. The "500 million years" figure is the authors' **estimate** from GFP diversification rates, not a measurement.
- [29] Jayatunga, M. K. P., Ayers, M., Bruens, L., Jayanth, D., & Meier, C. (2024). *How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons*. Drug Discovery Today 29(6), 104009. https://doi.org/10.1016/j.drudis.2024.104009 — accessed 2026-07-29. [type: peer-reviewed] — "In Phase I we find AI-discovered molecules have an 80–90% success rate, substantially higher than historic industry averages… In Phase II the success rate is ~40%, albeit on a limited sample size, comparable to historic industry averages." The Phase I figure rests on **24 molecules with 21 successes**; the authors describe the result as "early signs of potential". Read alongside [68].
- [30] Insilico Medicine et al. (2025). *A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial*. Nature Medicine, 3 Jun 2025. https://doi.org/10.1038/s41591-025-03743-2 — accessed 2026-07-29. [type: peer-reviewed] — The strongest positive clinical datapoint (n=71 across four arms) **and** its own main-text statement that "AI-discovered drugs have experienced similar levels of phase 2 trial failure as non-AI-discovered drugs, and none has so far progressed through phase 3 trials". **Phase-3 registration re-checked 2026-09-02** via the ClinicalTrials.gov API (https://clinicaltrials.gov/api/v2/studies/NCT07687459): "Study Evaluation Rentosertib (INS018_055) Administered Orally in Patients With Idiopathic Pulmonary Fibrosis (IPF)", overall status **NOT_YET_RECRUITING**, Phase 3, enrolment 320, sponsor InSilico Medicine Hong Kong Limited, record last updated 2026-07-01. So the paper's sentence still holds: registered is not progressed.
- [31] Kedzierska, K. Z., Crawford, L., Amini, A. P., & Lu, A. X. (2025). *Zero-shot evaluation reveals limitations of single-cell foundation models*. Genome Biology 26, 101. https://doi.org/10.1186/s13059-025-03574-x — accessed 2026-07-29. [type: peer-reviewed] — **A 2025 evaluation of 2023–24 single-cell foundation models (Geneformer, scGPT)**; the slide names both models and the year rather than implying a claim about the 2026 crop. Across five human tissue datasets, Geneformer and scGPT underperform highly variable gene selection, scVI and Harmony on cell-type clustering and batch integration — and several evaluation datasets overlapped with pretraining data, so the reported gap is a best case.
### D. **Engineering**
- [32] McGreivy, N., & Hakim, A. (2024). *Weak baselines and reporting biases lead to overoptimism in machine learning for fluid-related partial differential equations*. Nature Machine Intelligence 6, 1256–1269. https://doi.org/10.1038/s42256-024-00897-5 · https://arxiv.org/abs/2407.07218 — accessed 2026-07-29. [type: peer-reviewed] — Of articles claiming to beat a standard numerical method, **79% (60 of 76) compared against a weak baseline**; outcome reporting bias and publication bias are "widespread"; the literature is "overoptimistic".
- [33] Krishnapriyan, A. S., Gholami, A., Zhe, S., Kirby, R. M., & Mahoney, M. W. (2021). *Characterizing possible failure modes in physics-informed neural networks*. NeurIPS 2021. https://arxiv.org/abs/2109.01050 — accessed 2026-07-29. [type: peer-reviewed] — PINNs "can learn good models for relatively trivial problems" but "easily fail to learn relevant physical phenomena for even slightly more complex problems"; the failure is in optimisation, not capacity.
- [34] Azizzadenesheli, K., Kovachki, N., Li, Z., et al. (2024). *Neural operators for accelerating scientific simulations and design*. Nature Reviews Physics 6, 320–328. https://doi.org/10.1038/s42254-024-00712-5 — accessed 2026-07-29. [type: peer-reviewed] — Source of the claim that neural operators can "augment, or even replace, existing numerical simulators in many applications… providing speedups of four to five orders of magnitude". **Attribution warning: this is a self-review** — the authors are the developers of the Fourier Neural Operator reviewing their own method, its supporting figures (45,000×, 26,000×, 700,000×) do not fall inside the stated range, and they are inference-versus-solver timings, i.e. exactly the comparison class [32] finds weakly baselined in 79% of cases. The slide attributes it as a self-review and ties it to [32] rather than presenting the two side by side.
- [35] Lam, R., Sanchez-Gonzalez, A., Willson, M., et al. (2023). *Learning skillful medium-range global weather forecasting* (GraphCast). Science 382, eadi2336. https://doi.org/10.1126/science.adi2336 · https://arxiv.org/abs/2212.12794 — accessed 2026-07-29. [type: peer-reviewed] — 10-day forecasts at 0.25° in under a minute; outperforms the most accurate operational deterministic system on 90% of 1,380 verification targets. Does **not** assimilate observations — it is initialised from a numerical-weather-prediction analysis.
- [36] ECMWF (2026). *AIFS Machine Learning data*; and Haiden, T., & Chevallier, M. (2026). *Forecast performance 2025*, ECMWF Newsletter 187. https://www.ecmwf.int/en/forecasts/dataset/aifs-machine-learning-data · https://www.ecmwf.int/en/newsletter/187/news/forecast-performance-2025 — accessed 2026-07-29. [type: primary-doc] — AIFS Single operational since 2025-02-25, AIFS ENS since 2025-07-01, both upgraded to v2 on 2026-05-12; medium-range error reductions against the physics-based IFS "typically of the order of 5–15%". The order-of-magnitude gain is in cost, not skill.
- [37] Ben Bouallègue, Z., Raoult, B., & Chantry, M. (2026). *Farewell to the external AI models*. ECMWF AIFS Blog, 11 May 2026. https://doi.org/10.21957/fa8ad01483 — accessed 2026-07-29. [type: primary-doc] — After the IFS Cycle 50r1 upgrade changed the initial conditions, the fine-tuned ML models (GraphCast, Aurora, AIFS v1.1) **lost skill**, while the physics model gained it; ECMWF stopped running Pangu-Weather, GraphCast, Aurora and FourCastNet in real time. The cleanest published demonstration of ML dataset shift in an operational setting.
- [38] Ohana, R., McCabe, M., Meyer, L., et al. (2024). *The Well: a large-scale collection of diverse physics simulations for machine learning*. NeurIPS 2024 Datasets & Benchmarks, 44989–45037. https://arxiv.org/abs/2412.00568 · https://polymathic-ai.org/the_well/ — accessed 2026-07-29. [type: peer-reviewed] — 15 TB across 16 simulation datasets with a unified PyTorch interface; the open benchmark that makes the fix for [32] practical. Its own maintainers state the shipped baselines "should not be considered as state-of-the-art".
- [39] Autodesk (2020; product page 2026). *Topology Optimization is not Generative Design* (Fusion 360 blog); *Generative design for manufacturing*. https://www.autodesk.com/products/fusion-360/blog/topology-optimization-is-not-generative-design/ · https://www.autodesk.com/solutions/generative-design/manufacturing — accessed 2026-07-29. [type: primary-doc] — The vendor's own vocabulary for what commercial generative design takes as input, quoted correctly: **Preserve Geometry** — "a type of geometry that you apply to a body you want to include in the final shape of your design" — and **Obstacle Geometry** — "a body you want to exclude from the final shape" — plus Structural Loads, Structural Constraints, materials and manufacturing methods. **No text prompt anywhere in that loop as of July 2026**, a claim the deck explicitly date-stamps because the vendor has announced prompt-driven CAD as forthcoming. **Re-check attempted 2026-09-02: autodesk.com returned HTTP 403 to our fetcher, so the claim keeps its July-2026 verification date and the slide, the handout and this entry all say so out loud.** This is the one date-stamped claim in Session 7 that deliberately still reads "July 2026"; it is a claim date, not a landscape stamp. (An earlier draft of this deck attributed the phrase "hold out areas, preserved areas" to Autodesk; that is not Autodesk's wording and has been corrected.)
- [40] Zoo (2026). *Frequently Asked Questions*; *Text-to-CAD*; *Design Studio* (**v1.4.4, page last updated 2026-08-28**). https://zoo.dev/docs/faq · https://text-to-cad.zoo.dev/ · https://zoo.dev/design-studio — accessed 2026-09-02. [type: primary-doc] — **The highest-churn source in this deck: it moved twice in six weeks.** At the 2026-07-29 check it was Design Studio v1.3.7 (page updated 2026-07-22) and Text-to-CAD was documented as returning a **STEP** file with "$10.00 worth of API calls per month for free". On the 2026-09-02 re-check: version **v1.4.4**, the conversational agent is branded **Zookeeper**, and the documented output has changed — CAD-generation results "can include editable **KCL** project files" while "producing formats such as STEP requires a separate modeling and export step". **The "$10.00 per month for free" figure could not be re-confirmed** — the free plan is now described in included usage minutes rendered from an unresolved template variable — so that figure is dropped from the demo script rather than carried forward. The two warnings the demo quotes are **unchanged and still live on the FAQ**: the model is "still experimental", "Some results may not be as good as you expect", and "Zookeeper can make mistakes - it may misunderstand intent, produce incorrect geometry, or suggest designs that aren't manufacturable or safe". Note the divergence worth teaching: the *product page* now uses confident marketing language ("production-ready CAD") while the *FAQ* keeps both warnings. Grounds the engineering demo.
- [41] Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., & Karri, R. (2022). *Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions*. IEEE Symposium on Security and Privacy 2022, 754–768. https://doi.org/10.1109/SP46214.2022.9833571 · https://arxiv.org/abs/2108.09293 — accessed 2026-07-29. [type: peer-reviewed] — **A 2022 measurement on 2022-era GitHub Copilot, and labelled as such wherever it appears** (slide 17, handout, demo §5). 89 scenarios, 1,689 generated programs, approximately 40% found vulnerable. Peer-reviewed evidence used in place of vendor security reports. Not re-measured on the 2026 model generation; the deck states that gap rather than implying the rate is current.
- [42] Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). *Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity*. **Preprint.** https://arxiv.org/abs/2507.09089 — accessed 2026-07-29. [type: peer-reviewed] — 16 experienced open-source developers, 246 tasks: AI made them **19% slower** (95% CI +2% to +39%) while they believed they were 20% faster. **The title's phrase "Early-2025 AI" is the vintage and is quoted on the slide as the vintage** — this is a historical baseline on the early-2025 model cohort, not a claim about 2026 assistants, and no replication on the current generation was found. Still a preprint on 2026-09-02 (no journal reference on the arXiv record). Carried forward from Session 5; labelled a preprint on the slide.
### E. **Physical sciences & mathematics**
- [43] Merchant, A., Batzner, S., Schoenholz, S. S., Aykol, M., Cheon, G., & Cubuk, E. D. (2023). *Scaling deep learning for materials discovery* (GNoME). Nature 624, 80–90. https://doi.org/10.1038/s41586-023-06735-9 — accessed 2026-07-29. [type: peer-reviewed] — The claim under dispute: 2.2 million predicted structures, ~380,000 predicted stable, described as "an order-of-magnitude expansion in stable materials known to humanity".
- [44] Cheetham, A. K., & Seshadri, R. (2024). *Artificial Intelligence Driving Materials Discovery? Perspective on the Article: Scaling Deep Learning for Materials Discovery*. Chemistry of Materials 36(8), 3490–3495. https://doi.org/10.1021/acs.chemmater.4c00643 · https://escholarship.org/uc/item/9qx9t3kz — accessed 2026-07-29. [type: peer-reviewed] — The published rebuttal: "finding scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility"; "we have yet to find any strikingly novel compounds"; "[the work] does not report any new materials but reports a list of proposed compounds… they cannot yet be regarded as materials".
- [45] Szymanski, N. J., Rendy, B., Fei, Y., et al. (2023). *An autonomous laboratory for the accelerated synthesis of inorganic materials* (A-Lab). Nature 624, 86–91. https://doi.org/10.1038/s41586-023-06734-w — accessed 2026-07-29. [type: peer-reviewed] — **Read only in its corrected form; see [69].** The title now says "inorganic materials" (originally "novel materials"), and the corrected abstract reads: "Over 17 days of continuous operation, the A-Lab realized **36 compounds from a set of 57 targets**." The widely repeated "41 of 58 targets" and "43 novel materials" are pre-correction figures and are **not** carried onto any slide in this deck.
- [46] Leeman, J., Liu, Y., Stiles, J., Lee, S. B., Bhatt, P., Schoop, L. M., & Palgrave, R. G. (2024). *Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis*. PRX Energy 3, 011002. https://doi.org/10.1103/PRXEnergy.3.011002 — accessed 2026-07-29. [type: peer-reviewed] — The published rebuttal: re-examining all 43 products, "these errors unfortunately lead to the conclusion that no new materials have been discovered in that work"; "two thirds of the claimed successful materials… are likely to be known compositionally disordered versions of the predicted ordered compounds"; and "automated Rietveld analysis of powder x-ray diffraction data is not yet reliable".
- [47] Riebesell, J., Goodall, R. E. A., Benner, P., et al. (2025). *A framework to evaluate machine learning crystal stability predictions* (Matbench Discovery). Nature Machine Intelligence 7, 836–847. https://doi.org/10.1038/s42256-025-01055-1 · https://matbench-discovery.materialsproject.org/ — accessed 2026-09-02. [type: peer-reviewed] — The live, open leaderboard that replaced the argument with a measurement, ranking models on crystal stability, geometry optimisation and thermal conductivity; the peer-reviewed paper's own hedged wording is that models have "advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials". **The crisper phrase "robust enough to deploy them as triaging steps" appears only in the project's GitHub README, not in the journal article** — where the deck uses it, it is attributed to the README. **Model-count policy, revised in this pass:** earlier drafts kept the count off the slides entirely (41 at initial review, 42 on 2026-07-29, 42 on 2026-08-08). Slide 19 and the handout now carry it **as an explicitly dated snapshot** — "2026-09-02: 42 models, led by EquiformerV3+DeNS-OAM (F1 0.931, κ_SRME 0.093, DAF 6.074), added 7 April 2026" — followed by the sentence that the line is stale the moment it is read. The count is doing pedagogical work (a live board versus a headline) precisely *because* it decays, so it is shown with its date rather than suppressed. Verified on a fresh fetch 2026-09-02; the leaderboard is live and any snapshot is stale within weeks.
- [48] NSF–DOE Vera C. Rubin Observatory (2026). *Rubin Observatory Launches Real-Time Alerts*, 25 Feb 2026; *Alerts and brokers*. https://rubinobservatory.org/news/first-alerts · https://rubinobservatory.org/for-scientists/data-products/alerts-and-brokers — accessed 2026-07-29. [type: primary-doc] — First scientific alerts issued the night of 24 Feb 2026 (800,000 that night), rising to up to seven million alerts per night; the brokers "use machine learning algorithms to filter, sort, and classify the alerts". **Note the source's own inconsistency:** the news release names nine brokers, while the data-products page says "Seven 'full-stream' brokers will receive and process the full LSST alert stream" plus "Two 'down-stream' brokers". The deck tracks the page's structure (seven full-stream plus two down-stream) rather than asserting a single total. ML as triage at a scale no human pipeline can touch.
- [49] Hubert, T., Masoom, R., Barekatain, M., et al. / Google DeepMind (2025). *Olympiad-level formal mathematical reasoning with reinforcement learning* (AlphaProof). Nature, 12 Nov 2025. https://doi.org/10.1038/s41586-025-09833-y — accessed 2026-07-29. [type: peer-reviewed] — Precisely what happened at IMO 2024: AlphaProof solved three of the five non-geometry problems (P1, P2, P6) **from statements manually formalised in Lean by experts**, each requiring 2–3 days of test-time RL; the geometry problem P4 was solved by AlphaGeometry 2; the two combinatorics problems were not solved. Combined score 28/42 — silver-medal range, one point below gold. The authors note the bespoke training scale "is likely beyond the reach of most academic research groups".
- [50] Lean community (2026). *Lean 4 Web*; *LeanSearch*; *Mathlib statistics*. https://live.lean-lang.org/ · https://leansearch.net/ · https://leanprover-community.github.io/mathlib_stats.html · https://github.com/leanprover-community/lean4web — accessed 2026-07-29. [type: primary-doc] — The browser-based Lean 4 environment used in the mathematics demo; its maintainers scope it explicitly to "smallish" snippets and state that "serious Lean code development and larger projects are considered out-of-scope". LeanSearch provides semantic search over Mathlib4. The Mathlib statistics page reported **283,218 theorems, 134,717 definitions and 772 contributors** on 2026-07-29 and **286,514 theorems, 136,111 definitions and 772 contributors** on a re-fetch **2026-09-02** — the counts advance daily, which is why the slide and the demo script both carry the date beside the number.
- [51] Epoch AI (2026). *FrontierMath*; *FrontierMath Tiers 1–4*. https://epoch.ai/frontiermath · https://epoch.ai/benchmarks/frontiermath-tier-4-v2 — accessed 2026-07-29. [type: primary-doc] — On 2026-06-12 Epoch released v2 "addressing errors in 42% of problems" (123 corrected in Tiers 1–3, 12 in Tier 4; 5 and 7 removed); the dataset now holds 338 problems, 12 of them public. The benchmark-hygiene lesson: a headline score is only as good as the benchmark's last audit.
- [52] IMO 2025 Organisers / Dolinar, G. (2025). *The 66th International Mathematical Olympiad draws to a close today*, 19 July 2025. https://imo2025.au/news/the-66th-international-mathematical-olympiad-draws-to-a-close-today/ — accessed 2026-07-29. [type: primary-doc] — The IMO's own statement about the AI systems that ran on its 2025 problems: "for the first time, a selection of AI companies were invited to join a **fringe event**… These companies also **privately tested closed-source AI models** on this year's problems". And, verbatim from the IMO President: "the IMO **cannot validate the methods**, including the amount of compute used or whether there was any human involvement, or whether the results can be **reproduced**. What we can say is that correct mathematical proofs, whether produced by the brightest students or AI models, are valid." Also records the human baseline: 641 students from 112 countries, five perfect scores. **2026-cycle check (2026-09-02):** imo-official.org confirms the 2026 olympiad was held in Shanghai, 10–21 July 2026, but carries no comparable statement about AI participation or validation, and no primary IMO source for a 2026 AI claim could be verified — press and prediction-market coverage exists and is **not** citable under this session's evidence rule. The deck therefore keeps the 2025 statement and marks it "as of the 2025 cycle" rather than extending the story.
### F. **Social sciences & humanities**
- [53] Gilardi, F., Alizadeh, M., & Kubli, M. (2023). *ChatGPT outperforms crowd workers for text-annotation tasks*. PNAS 120(30), e2305016120. https://doi.org/10.1073/pnas.2305016120 — accessed 2026-07-29. [type: peer-reviewed] — **Measured in 2023 on ChatGPT (GPT-3.5-era); the slide and handout now say "ChatGPT in 2023" rather than leaving the model implicit.** n = 6,183 texts across four datasets: zero-shot accuracy exceeded **MTurk crowd workers** — not trained annotators — by about 25 percentage points on average, at under $0.003 per annotation. **Two qualifications the headline drops:** absolute accuracy was a modest **59–83%**, and the result about trained annotators concerns **intercoder agreement**, not accuracy. Says nothing about interpretive coding, non-English text, or downstream inference validity.
- [54] Egami, N., Hinck, M., Stewart, B. M., & Wei, H. (2023). *Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models*. NeurIPS 2023. https://arxiv.org/abs/2306.04746 · http://naokiegami.com/dsl/ — accessed 2026-07-29. [type: peer-reviewed] — The methodological load-bearing source: direct use of surrogate labels in downstream analyses "leads to substantial bias and invalid confidence intervals, **even with high surrogate accuracy of 80–90%**". The fix requires the researcher to control the probability of sampling documents for gold-standard human labelling — a design decision made before annotating.
- [55] Ollion, É., Shen, R., Macanovic, A., & Chatelain, A. (2024). *The dangers of using proprietary LLMs for research*. Nature Machine Intelligence 6, 4–5. https://doi.org/10.1038/s42256-023-00783-6 — accessed 2026-07-29. [type: peer-reviewed] — "an experiment carried out with a given model often yielded different results when repeated a few weeks later"; model deprecation makes "reproducibility nearly impossible"; plus the corpus-licensing constraint that forced one large study onto article titles only.
- [56] Liu, X., Zambrano, A. F., Baker, R. S., et al. (2025). *Qualitative Coding with GPT-4: Where it Works Better*. Journal of Learning Analytics 12(1), 169–185. https://doi.org/10.18608/jla.2025.8575 — accessed 2026-07-29. [type: peer-reviewed] — κ ≥ 0.70 for 25 of 34 constructs, but only when the best prompting strategy is chosen per construct — "No single method consistently outperforms the others". Zero-shot ranged from κ = 0.91 (*Questioning*) to κ = 0.11 (*Direct Instruction*). The finding to put on a slide, quoted exactly: "GPT-4 has the most difficulty with the same constructs than human coders find more difficult to reach inter-rater reliability on." Human–human κ itself ranged 0.24–0.87. (Earlier drafts paraphrased this inside quotation marks; corrected.) **Model naming fixed in this pass:** slide 23 previously called this "a frontier model", which in 2026 reads as a claim about the current generation. It is **GPT-4**, measured in a 2025 paper, and the slide, the handout and demo §7 all now name it and add that a better model raises the κ without making an ambiguous construct unambiguous. The κ range is not re-measurable by us on current models, so it stays as a labelled 2023-model baseline.
- [57] Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C., & Wingate, D. (2023). *Out of One, Many: Using Language Models to Simulate Human Samples*. Political Analysis 31(3), 337–351. https://doi.org/10.1017/pan.2023.2 — accessed 2026-07-29. [type: peer-reviewed] — **GPT-3, published 2023.** The origin of "silicon samples" and "algorithmic fidelity"; the pro side of the contested story. Slide 24 now carries the model and year in the column itself, and frames the whole exchange as the historical record of a debate fought on 2022–24 models that nobody has re-run on the 2026 generation.
- [58] Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). *Synthetic Replacements for Human Survey Data? The Perils of Large Language Models*. Political Analysis 32(4), 401–416. https://doi.org/10.1017/pan.2024.5 — accessed 2026-07-29. [type: peer-reviewed] — **GPT-3.5, published 2024.** Means match the baseline survey, but "there is less variation in responses than in the real surveys, and regression coefficients often differ significantly"; distributions shift with minor prompt-wording changes, and "the same prompt yields significantly different results over a 3-month period".
- [59] Boelaert, J., Coavoux, S., Ollion, É., Petev, I., & Präg, P. (2025). *Machine Bias: How Do Generative Language Models Answer Opinion Polls?* Sociological Methods & Research. https://doi.org/10.1177/00491241251330582 — accessed 2026-07-29. [type: peer-reviewed] — **Published 2025.** Independent second rebuttal: "models cannot replace research subjects for opinion or attitudinal research"; they show "a strong bias and a low variance on each topic", and "this bias randomly varies from one topic to the next" — so it cannot be corrected without the human data you were trying to avoid collecting.
- [60] Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., & Bernstein, M. S. (2026). *LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals*. **Preprint, v3, last revised 2026-06-28; never journal-published — status re-checked 2026-09-02: still v3, still a preprint, no journal reference or DOI beyond the arXiv DOI.** https://arxiv.org/abs/2411.10109 — accessed 2026-09-02. [type: peer-reviewed] — **This source was retitled and its numbers revised.** The widely quoted "Generative Agent Simulations of 1,000 People" title and the "85%" figure exist only in v1. Current wording: across 1,052 Americans, "interview-only, survey-only, and combined agents achieved accuracies equal to **83%, 82%, and 86%** of participants' own two-week test-retest consistency benchmark, respectively, compared with **74%** for demographics-only agents", with "gains over either source alone were **modest**". The revision undercuts the interview-depth story: a structured survey buys within one point of a two-hour interview. Labelled an unpublished preprint on the slide.
- [61] Vesuvius Challenge (2026). *An entire Herculaneum scroll has been read for the first time*, 25 Jun 2026. https://scrollprize.org/firstscroll · https://scrollprize.org/data — accessed 2026-07-29. [type: primary-doc] — PHerc. 1667 virtually unwrapped and read end to end: roughly 1.4 m of papyrus and around twenty-two columns of Greek, from X-ray tomography plus geometric reconstruction plus ML ink detection, every reading transcribed and reviewed by papyrologists. Their own caveat: "Because the papyrus is damaged, the readings are fragmentary, with gaps where the surface is lost." Data openly licensed, code on GitHub, preprint posted.
- [62] READ-COOP SCE / Transkribus (2026). *Pricing*; *Character Error Rate (CER) — The Standard Metric for Transcription Accuracy*. https://www.transkribus.org/pricing · https://www.transkribus.org/character-error-rate-cer-explained — accessed 2026-07-29. [type: primary-doc] — Free plan: 50 credits per month, one seat; Scholar €99/year. The vendor's own indicative CER bands: <2% excellent, 2–5% good, 5–10% needs review; printed post-1950 0.5–1%, 19th-century handwriting 2–5%, medieval manuscripts 5–15%; **15–30 pages** of hand-corrected ground truth to train a model on a **single hand**, and **50–100 pages** for a **multi-hand collection** — two distinct requirements that must not be collapsed into one range. Vendor-stated, not independently evaluated.
- [63] Hutchinson, D. (2026). *Benchmarking as Source Criticism: From Recognition to Reasoning in LLM Assessment*. Journal of Open Humanities Data 12, art. 43. https://doi.org/10.5334/johd.489 — accessed 2026-07-29. [type: peer-reviewed] — **Synthesises five existing benchmarks** of LLM historical competence rather than reporting original experiments, and on that basis reports "a collapse in LLM performance as assessments move from contaminated to decontaminated datasets, from Western to global knowledge domains, and from multiple-choice questions to open-ended responses".
- [64] MAXQDA/VERBI, Lumivero (NVivo), ATLAS.ti and Taguette (2026). *AI feature documentation*. https://www.maxqda.com/products/ai-assist · https://lumivero.com/products/nvivo/ · https://atlasti.com/ · https://www.taguette.org/about.html — accessed 2026-07-29. [type: primary-doc] — What the QDA vendors say their AI features do: coding *suggestions* the researcher accepts or rejects, code-label suggestions, document and code summaries, chat-with-your-data with reference markers back to source passages. None claims to produce validated codes; none publishes an independent inter-rater reliability evaluation of its AI coder. Taguette is free, open source and ships **no** AI features — the deliberate contrast for IP-restricted or sensitive material.
### G. Tracking the frontier — venues and benchmarks named on the Field-Frontier Tracker
- [65] NEJM Group (2026). *About NEJM AI*. https://ai.nejm.org/about — accessed 2026-07-29. [type: primary-doc] — Monthly, digital-only journal (ISSN 2836-9386) that "intentionally pairs 'pre-clinical' and clinical articles"; the venue where medicine's AI evidence is now appraised rather than announced.
- [66] Protein Structure Prediction Center (2026). *CASP16*. https://predictioncenter.org/casp16/ — accessed 2026-07-29. [type: primary-doc] — The independent, blinded community assessment that decides what structure-prediction methods can actually do; "tens of thousands of models submitted by approximately 100 research groups worldwide", evaluated by independent assessors as experimental coordinates become available.
- [67] Summer Institutes in Computational Social Science (2026). *About SICSS*. https://sicss.io/about — accessed 2026-07-29. [type: primary-doc] — Free training in computational social science, founded 2017 by Chris Bail and Matthew Salganik; the standing entry point for social scientists tracking method change.
### H. Added during independent verification (2026-07-29; [70] added 2026-08-08 during the fix pass)
These entries were added after a primary-source re-check of the field claims. They are numbered in append order so that no existing `[n]` marker had to move; the sections above are otherwise in slide order.
- [68] Jacobson, R. D. (2025). *The AI drug revolution needs a revolution*. npj Drug Discovery, 4 June 2025. https://doi.org/10.1038/s44386-025-00013-6 — accessed 2026-07-29. [type: peer-reviewed] — A peer-reviewed challenge to the framing of AI drug discovery itself: that the field optimises target and molecule discovery "in a human-agnostic manner" while leaving untouched the measurement of human response that actually kills candidates in the clinic. Balances the Phase I success figure in [29] with a critique of what that figure does and does not predict.
- [69] Szymanski, N. J., Rendy, B., Fei, Y., et al. (2026). *Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials*. **Nature 650, E1** (published online 19 January 2026; issue of 5 February 2026). https://doi.org/10.1038/s41586-025-09992-y — accessed 2026-07-29. [type: peer-reviewed] — Read directly rather than through any summary, because two independent summaries of it disagreed. It states three distinct things. **(i)** The novelty concession: "the original claims of material novelty were subject to misinterpretation—their intention was to indicate that the materials were **new to the prediction platform, not necessarily new to science**." **(ii)** The re-analysis: "we have manually re-analyzed the diffraction patterns and have confirmed that the prediction platform came to the correct conclusion in **36 of its 40 reported successes, with 4 compounds being inconclusive**", peer-reviewed post-publication; plus one compound (Zn2Cr3FeO8) removed as training-data contamination. **(iii)** Consequently the corrected article abstract now reads "**36 compounds from a set of 57 targets**". Those two denominators are consistent, not contradictory: 40 reported successes minus 4 inconclusive = 36, against 58 targets minus the removed compound = 57. Nature "thanks the correspondents who brought these issues to our and the authors' attention". **This entry is the single strongest piece of evidence on slide 19, because the slide's thesis is that corrections to AI-for-science claims arrive in print.**
- [70] Kim, D., Woodbury, S. M., Ahern, W., et al. (2025). *Computational design of metallohydrolases*. Nature 649, 246–253 (published online 2025-12-03). https://doi.org/10.1038/s41586-025-09746-w — accessed 2026-08-08. [type: peer-reviewed] — Read directly (PMC full text) because [27] deliberately carried no k_cat/K_M number pending this check, and the number a prior draft of this deck had used (16,000 ± 2,000 M⁻¹s⁻¹) needed verifying before being trusted. This is the RFdiffusion2 enzyme-scaffolding paper's own accompanying paper. **First round of 96 designs**, theozymes without an explicit catalytic base: best design (ZETA_1), verbatim "a k_cat/K_M of 16,000 ± 2,000 M⁻¹ s⁻¹ for A1, the most active design" — a real figure, but not the paper's best. **Second round of 96 designs**, theozymes now built with an explicit general base to activate the water: the paper's best result overall, ZETA_2, verbatim "k_cat/K_M = 53,000 ± 5,000 M⁻¹ s⁻¹" with "k_cat = 1.5 ± 0.1 s⁻¹"; two further second-round designs reached 19,000 ± 2,000 and 1,100 ± 200 M⁻¹ s⁻¹, and five of the second-round designs exceeded 10⁴ M⁻¹ s⁻¹. **This is the figure carried onto slide 13**, attributed to the second 96-design set and its general base, not presented as a first-round or zero-shot number.
---
<!--
Deliberately NOT cited, and why — recorded so the next author does not re-add them:
* ic2s2.org — the former IC2S2 conference domain now resolves to unrelated commercial
spam content. Do not link it. Journals ([58], [59]) and SICSS ([67]) are the stable
entry points instead.
* Vendor security "GenAI code" reports — vendor research, no external replication.
Replaced by the peer-reviewed [41].
* Any FDA device count sourced to trade press. [7] is FDA's own list and file.
* A PoseBusters percentage for AlphaFold 3 — the figure lives in an Extended Data table
that could not be read directly. [19]'s P-values and stated limitations are used instead.
* "173 AI drug programs in clinical trials" and similar aggregate pipeline counts —
traceable only to content farms. [29] and [30] carry the pipeline claim instead.
* Adoption/reach figures for clinical Q&A products — company marketing, unverified.
-->
---
*AI for Researchers · Session 7: Discipline Deep-Dives: Specialized AI by Field · Landscape as of August 2026*