AI for Researchers · Visual Deck

Session 7: Discipline Deep-Dives: Specialized AI by Field

Five fields, five landmark systems, and the one question that tells you whether any of them is safe to use

{{Presenter Name}}

Landscape as of August 2026

AI for ResearchersLandscape as of August 20261 / 33

2-Minute Recap · The Whole Series

The Series So Far


flowchart LR
  s1(["1 · Map · Taxonomy
· CRIT"]):::rest --> s2(["2 · Tool
Matrix"]):::rest s2 --> s3(["3 · Summary-
Trust Triage"]):::rest s3 --> s4(["4 · Policy
Table"]):::rest s4 --> s5(["5 · Canvas —
row 7 empty"]):::rest s5 --> s6(["6 · Orientation
Ladder"]):::rest s6 --> s7(["7 · Field-Frontier
Tracker

today"]):::today classDef today fill:#ab7d22,stroke:#d9b36c,color:#14132b,font-size:19px classDef rest fill:#23204c,stroke:#8f8cb8,color:#b9b7d6,font-size:19px

Today (Session 7): the AI that is not a chatbot — the systems reshaping five disciplines, what each one really does, and how to track yours after this series ends.
Row 7 of Session 5's canvas gets filled today, and the canvas is finished.

AI for ResearchersLandscape as of August 20262 / 33

Learning Objectives

By the End of This Session You Will Be Able To…


1Identify the specialized, non-chatbot AI tools transforming your discipline
2Explain one specialized application in each of five fieldsmedicine · life sciences · engineering · physical sciences/math · social sciences/humanities
3Evaluate a field-specific AI tool relevant to your own research area
4Track your field's AI frontierkey venues, benchmarks, and review articles
5Compare specialized tools to the general-purpose tools from earlier in the series

Full verbatim wording in the reference deck and curriculum

AI for ResearchersLandscape as of August 20263 / 33

Section 01

The Tier the Taxonomy Left Out

Six sessions on tools that write. This one is about tools that predict a structure, propose a compound, or refuse a proof — and why they fail in a completely different direction.

AI for ResearchersLandscape as of August 20264 / 33

Literacy Foundation — The Key Slide

Specialized AI Is Not a Better Chatbot. It Is a Different Object.


flowchart LR
  M["a domain model
not a chat box"]:::model --> O["a domain object, not prose
a coordinate file · a candidate crystal ·
a proof term · a coded transcript"]:::obj --> V["a verifier that is not software
a wet lab · a telescope ·
a proof kernel · a second human coder"]:::ver classDef model fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef obj fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef ver fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px
decisive before the chatbots: the 2024 Nobel Prizes in Physics [2] and Chemistry [1]
scored by somebody else — blinded, external, repeated: CASP [66], Matbench Discovery [47], an FDA premarket review [7]
you cannot judge it by reading it — that verifier is the whole safety story [3]

Tier 1/2/3 sorted tools by how general they are. Today's question is different and better: what checks it, and did you run the check?

AI for ResearchersLandscape as of August 20265 / 33

Series Artefact — The Reusable Part of Today

Five Questions to Ask Any Field-Specific AI Tool


1What does it output — and what does its own paper say it does not do?AlphaFold 3's limitations section names a 4.4% chirality violation rate [19]
2Who scored it — and were they independent of the builders?a vendor benchmark and a blinded community assessment are not the same evidence [66] [47]
3What is the gap between the benchmark and the deployment?trial-matching: 87.3% on synthetic patients [16] — precision 0.33 on 157 real ones [17]
4What is the external verifier — and can you afford to run it?ninety-six wells · a telescope night · a proof kernel · a second coder
5What happens when the inputs shift?when ECMWF upgraded its physics model, the fine-tuned ML forecasters lost skill [37]

The spine of the Field-Frontier Tracker (slide 26). Write them on the inside cover of your notebook; the tool names will change by Christmas.

AI for ResearchersLandscape as of August 20266 / 33

Section 02

Medicine & Clinical Research

The only field on today's tour where a regulator already decides what the software is allowed to claim — and where the published numbers fall furthest between the benchmark and the clinic.

AI for ResearchersLandscape as of August 20267 / 33

Medicine — Landmark Systems, Regulatory Context

1,524 Authorised Devices, and Almost None of Them Is What You Picture


1,524entries counted by us in the 2026-06-16 download of FDA's AI-Enabled Medical Device List — decisions 1995 to 30 March 2026 [7]
76.4%have Radiology as lead review panel — our count, same download; FDA publishes none of these numbers as prose [7]
0identified as large-language-model based — and still 0 on the 2026-09-02 re-check; FDA "will explore" tagging foundation models, future tense [7]

The autonomous exemplar is narrower than its reputation: IDx-DR — one disease, one camera, one referral decision; "not intended to detect concomitant diseases" [8]. And FDA's own words: the list is "not a comprehensive resource" — those counts are a snapshot [7].

AI for ResearchersLandscape as of August 20268 / 33

Medicine — What Changed for Practitioners

The Number That Travels, and the Number That Does Not


87.3% → 0.33Trial matching, 2024–25 systems. TrialGPT: 87.3% criterion accuracy on 183 synthetic patients [16]; four deployed tools on 157 real tumour-board patients: mean precision 0.33, recall 0.32, 38% got no trial at all [17]
66%, but 31%Imaging. 173 CE-certified radiology products: any peer-reviewed evidence rose 36% → 66% — yet only 31% at the level of decisions, outcomes or cost; prospective designs 19% → 16% [11]
87.2% / 90.7%Where evidence is strong, it is narrow and prospective. Autonomous DR screening: pivotal-trial sensitivity / specificity, 10 US primary-care sites [9]; pooled over 887,244 examinations: 0.93 / 0.90 [10]

No comparable real-patient evaluation of a 2026-generation matcher was found — the gap, not the number, is what transfers. The meta-analysis itself calls for "post-market audits with standardized gradability metrics" [10] [12].

AI for ResearchersLandscape as of August 20269 / 33

Medicine — The Caution

Twenty Hours of AI Training Did Not Protect Them


84.9% → 73.3%composite diagnostic accuracy when exposed to erroneous LLM recommendations — RCT run in 2025 on a GPT-4-class assistant; 44 physicians, 264 cases [15]
−14.0 ptsadjusted (95% CI −18.9 to −9.1, P < 0.0001) — after a 20-hour AI-literacy course, with voluntary consultation [15]
2 pts · 6.5 ptsthe companion RCTs, 2023–25 cohort: +2 (P = .60) in 2024 [13]; +6.5 in a 2025 GPT-4 trial where the model alone was indistinguishable from model-plus-physician (−0.9%, P = 0.8) [14]

No RCT of a 2026-generation assistant has replaced these — plan for automation bias, a property of use, not of a model version. The least conversational tools work best: PubMed-wide entity search, F-score 82.0%, shows you the sentence [18]. Disclosure for your paper: Session 4's Policy Table.

AI for ResearchersLandscape as of August 202610 / 33

Section 03

Biotechnology & Life Sciences

The field with the clearest win of the decade, the most disciplined published limitations, and the most contested commercial promise.

AI for ResearchersLandscape as of August 202611 / 33

Biotech — Landmark System, From Its Own Paper

AlphaFold 3: What It Does, and What Its Authors Say It Does Not


Does

  • Predicts the joint structure of complexes — proteins, nucleic acids, small molecules, ions, modified residues [19]
  • Beats classical docking without structural inputs (vs Vina, P = 2.27 × 10⁻¹³) [19]
  • Underwrites a database of over 200 million predicted structures, free [22]
  • Blinded verdict, CASP16: single-domain fold prediction "nearly solved" [21]

Does not

  • Give dynamics: static outputs; multiple seeds "do not approximate the solution ensemble" [19]
  • Guarantee chemistry: 4.4% chirality violation rate; hallucinated order in disordered regions [19]
  • Rank outputs: "model ranking remains a persistent weakness" [21]
  • Come without strings: 30 jobs/day, closed ligand list, 2021-09-30 cutoff default; weights gated, non-commercial [20]

Every line from the AF3 paper [19], DeepMind's own docs [20], or the independent CASP16 assessment [21] [23] — never coverage. An AF3 model that looks clean is not evidence that it is.

AI for ResearchersLandscape as of August 202612 / 33

Biotech — What Changed for Practitioners

Design Moved From Reading Proteins to Writing Them — at a Measured Cost


19%RFdiffusion binders: experimental success, 95 designs tested per target across five targets — "roughly two orders of magnitude" over Rosetta [25]; BindCraft: 10–100% across twelve targets, averaging 46.3% [26]
53,000 ± 5,000 M⁻¹s⁻¹best RFdiffusion2 enzyme design — actives found "testing fewer than 96 sequences"; its 41-of-41 site benchmark is in-silico [27] [70]

The design software is free; the verifier is your consumables budget — question 4. "41 of 41" is computational; "19%" and "46.3%" are experimental. Never compare across that line. ESM3 generates sequence, structure and function jointly [28].

AI for ResearchersLandscape as of August 202613 / 33

Biotech — Where the Story Is Contested

AI Drug Discovery: State the Claim, Then State Its Own Authors' Caveat


"In Phase I… an 80–90% success rate… In Phase II… ~40%, albeit on a limited sample size, comparable to historic industry averages."

Drug Discovery Today 2024 — on 24 molecules, 21 successes; the authors call this "early signs of potential", not proof [29]

"AI-discovered drugs have experienced similar levels of phase 2 trial failure… and none has so far progressed through phase 3 trials."

Printed inside the most positive clinical paper of 2025 (phase 2a, n = 71: FVC +98.4 mL top dose vs −20.3 mL placebo). First registered phase 3, NCT07687459, still "not yet recruiting" on 2026-09-02 [30]

A peer-reviewed challenge to the framing itself [68] — and one layer down: Geneformer and scGPT, zero-shot, underperformed highly variable gene selection, scVI and Harmony [31]. A 2025 result on 2023–24 models: re-run it before you cite it.

AI for ResearchersLandscape as of August 202614 / 33

Section 04

Engineering

One operational triumph, one measured literature-wide bias, and a word — "generative design" — that does not mean what the audience assumes it means.

AI for ResearchersLandscape as of August 202615 / 33

Engineering — Landmark System, and the Published Correction

Simulation Surrogates: Real, Operational, and Systematically Oversold


25 Feb 2025ECMWF's ML forecast system AIFS Single goes operational; the ensemble follows 1 July 2025; both to v2 on 12 May 2026 [36]
5–15%the honest accuracy gain vs the physics-based IFS — the order-of-magnitude win is in cost, not skill [36]
79%of ML-for-fluid-PDE papers claiming to beat a numerical method compared against a weak baseline (60 of 76) [32]

The "four to five orders of magnitude" speedup review was written by the method's own developers [34]; PINNs "easily fail… for even slightly more complex problems" — an optimisation failure [33]. Treat any order-of-magnitude speedup as unverified until you see the baseline [32].

AI for ResearchersLandscape as of August 202616 / 33

Engineering — What Changed, and the Caution

"Generative Design" Is Not a Text Prompt — and Your Model Ages


verified July 2026 — re-date this oneNo text prompt in the loop: inputs are preserve geometry, obstacle geometry, loads, constraints, materials, manufacturing methods [39] — the vendor blocked our 2026-09-02 re-check and has announced prompt-driven CAD
v1.4.4 · Aug 2026Text-to-CAD is separate and younger: "still experimental", designs that "aren't manufacturable or safe" [40] — a first solid, not an assembly
"Early-2025 AI"19% slower (95% CI +2% to +39%): 16 experienced developers on their own repos, while believing they were 20% faster [42]
2022-era Copilot~40% of 1,689 programs vulnerable across 89 security scenarios [41] — historical measurements on named generations: quote them with the year, or not at all

Dataset shift is not hypothetical: when ECMWF upgraded its physics, the fine-tuned forecasters lost skill and ECMWF stopped running them [37] — why open benchmarks exist [38]. [42] is still a preprint.

AI for ResearchersLandscape as of August 202617 / 33

Section 05

Physical Sciences & Mathematics

Where two disciplines built the best answer anyone has to AI hype: they let outsiders re-examine the claim. One of them got a correction. The other could not even try.

AI for ResearchersLandscape as of August 202618 / 33

Physical Sciences — Where the Story Is Contested

Materials Discovery: the Claim, and the Peer-Reviewed Rebuttal


2.2M · 380kThe claim (Nature, 2023). GNoME structures predicted / ~stable — "an order-of-magnitude expansion in stable materials known to humanity" [43]; a companion lab reported novel materials over 17 days [45]
0 · 0The rebuttals (peer-reviewed, 2024). "Scant evidence… yet to find any strikingly novel compounds" [44]; all 43 A-Lab products re-examined: "no new materials have been discovered in that work" [46] — yet both call the method "sound" [44]
36 of 57The Author Correction (Nature, 2026) restates the result — "new to the prediction platform, not necessarily new to science" [69]

The live answer is a leaderboard, not a headline: Matbench Discovery — models "effectively and cheaply pre-screen" candidates [47]. Snapshot, 2026-09-02: 42 models, led by EquiformerV3+DeNS-OAM (F1 0.931), added April 2026 — stale the moment it is read.

AI for ResearchersLandscape as of August 202619 / 33

Physical Sciences — What Changed for Practitioners

In Observational Physics, AI Is Not the Discovery. It Is the Only Way to Get to It.


flowchart LR
  SKY["Rubin Observatory
800,000 alerts on night one —
up to seven million per night"]:::io --> BR["brokers — seven full-stream,
two down-stream: ML filter,
sort and classify"]:::model --> CAND["candidates
public alert ~two minutes
from the shutter"]:::hot --> FU["follow-up observation
= the evidence"]:::ver classDef io fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef model fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef hot fill:#23204c,stroke:#d9b36c,color:#eceaf8,font-size:19px classDef ver fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px

The model produces candidates; the instrument produces evidence — where the two get confused, the correction takes years [44] [46].
Rubin's alerts are world-public [48] — ask who is allowed to re-examine your field's classifier. Very few AI claims survive that test.

AI for ResearchersLandscape as of August 202620 / 33

Mathematics — Landmark System, Stated Exactly

Theorem Provers: the Only Field Whose Verifier Cannot Be Fooled


flowchart LR
  CB["a chatbot's fluent proof
true or false — you cannot tell by reading it"]:::n --> L["formalised
in Lean"]:::n --> K{"proof
kernel"}:::kern K -- accepts --> OK["machine-checked — the one AI output
you do not have to trust"]:::ok K -- rejects --> NO["rejected —
however good the prose"]:::bad classDef n fill:#23204c,stroke:#8f8cb8,color:#eceaf8,font-size:19px classDef kern fill:#1c1a3f,stroke:#d9b36c,color:#d9b36c,font-size:19px classDef ok fill:#1c1a3f,stroke:#2fae74,color:#7fd6a4,font-size:19px classDef bad fill:#1c1a3f,stroke:#c14b58,color:#ef9a9a,font-size:19px
3 of 5non-geometry IMO 2024 problems AlphaProof proved in Lean — from statements manually formalised by humans, 2–3 days of test-time RL each; both combinatorics problems unsolved; combined 28/42 [49]
286,514theorems in Lean's Mathlib, 772 contributors (2026-09-02; it advances daily) [50]

The contrast: July 2025's gold-medal claims came from undisclosed systems — the IMO "cannot validate the methods… or whether the results can be reproduced" [52], still its last word. And one research-maths benchmark shipped a v2 "addressing errors in 42% of problems" [51].

AI for ResearchersLandscape as of August 202621 / 33

Section 06

Social Sciences & Humanities

The field where the tool is a general model doing a specialised job — and where accuracy on the label is the easy half of the problem.

AI for ResearchersLandscape as of August 202622 / 33

Social Sciences — Landmark Result, and the Correction

Annotation at Scale Works. Plugging the Labels Into a Regression Does Not.


25 of 34constructs reached κ ≥ 0.70 with GPT-4 — but only with the best prompting strategy chosen per construct; "no single method consistently outperforms the others" [56]
80–90% is not enoughsurrogate labels in downstream analyses: "substantial bias and invalid confidence intervals, even with high surrogate accuracy of 80–90%" — errors are non-random [54]

ChatGPT in 2023 beat MTurk crowd workers by ~25 points at under $0.003 — but absolute accuracy was 59–83% [53]. Manifest codes transfer; interpretive ones do not [56] — and repeated runs drift, so record model version and date, every time [55].

AI for ResearchersLandscape as of August 202623 / 33

Social Sciences — Where the Story Is Contested

"Silicon Samples": a Debate Fought on 2022–24 Models — and Still Unsettled


The claim

  • GPT-3, 2023. Can "accurately emulate response distributions from a wide variety of human subgroups" — the authors' term: algorithmic fidelity [57]
  • Preprint, revised June 2026. Across 1,052 Americans: interview-, survey- and combined agents at 83%, 82% and 86% of participants' own test–retest consistency, vs 74% demographics-only [60]

The rebuttals

  • GPT-3.5, 2024. "Less variation in responses than in the real surveys"; regression coefficients often differ; the same prompt drifts over a 3-month period [58]
  • Independent replication, 2025. Models "cannot replace research subjects": strong bias, low variance — and the bias "randomly varies from one topic to the next" [59]

A live argument, not a verdict: every number is from a 2022–24 model, and nobody has re-run either side on the 2026 generation. Matching the mean is not matching the distribution — and [60] is still an unpublished preprint (v3, checked 2026-09-02): a structured survey (82%) buys what a two-hour interview (83%) does.

AI for ResearchersLandscape as of August 202624 / 33

Humanities — The Best-Verified Result in This Deck

A Scroll Nobody Could Open, Read End to End


"PHerc. 1667, sealed since the eruption of Vesuvius in 79 AD, has been virtually unwrapped and read from beginning to end… roughly 1.4 metres of papyrus and around twenty-two columns of Greek."

Vesuvius Challenge, 25 June 2026 — X-ray tomography, geometric reconstruction, ML ink detection; every reading reviewed by papyrologists; data openly licensed, code on GitHub; the announcement names its own limit: readings "fragmentary, with gaps where the surface is lost" [61]

<2% · 2–5% · 5–15%the everyday tool is handwritten text recognition — vendor's character-error-rate bands: publication-ready / good / medieval manuscripts [62]
15–30 pagesof hand-corrected ground truth for a single hand; 50–100 for a multi-hand collection — vendor-stated, not independently evaluated [62]

The caution: "a collapse in LLM performance" from contaminated to decontaminated datasets, Western to global domains, multiple-choice to open-ended [63]. Recognition is not reasoning.

AI for ResearchersLandscape as of August 202625 / 33

Series Artefact · Today's Canonical Workflow

The Field-Frontier Tracker


Field
The external verifier
Where the frontier is published
Medicineclinical research
Premarket review, then a prospective study in your own setting [10]
FDA's device list, updated in place [7] · NEJM AI [65] · your specialty's registry
Biotechlife sciences
CASP, blinded [66] — then 96 wells: binder success 19% [25] or 10–100% [26]
CASP assessment papers [66] [21] · AlphaFold DB release notes [22] · bioRxiv, then the journal
Engineering
A strong classical baseline on an open dataset — 79% of comparisons used a weak one [32]
Open benchmark suites such as The Well [38] · ECMWF's AIFS pages, which publish the failures too [36] [37]
Physical sciencesand mathematics
Synthesis and diffraction re-examined by outsiders — and a journal Author Correction [45] [46]; a live leaderboard [47]; a proof kernel [50]
Matbench Discovery [47] · Rubin broker docs [48] · Mathlib and the Lean Zulip [50] · benchmark changelogs [51]
Social sciencesand humanities
Human double-coding on a random subset, sampling probability fixed in advance [54]
Sociological Methods & Research · Political Analysis [58] [59] · J. Open Humanities Data [63] · SICSS [67]

The verifier and the venue are the whole artefact — a field's frontier is not a list of tools. Rebuild it for your own field; it stays true when every product name changes. Full five-column table in the reference deck.

AI for ResearchersLandscape as of August 202626 / 33

Series Artefact — Session 5's Canvas, Row 7

The Last Row. The Canvas Is Finished.


1–5Frame · Discover · Read · Analyse · WriteSessions 1–5 · Tier 1–2 · mostly Unreliable — as recorded on your canvas
6Cross-field — field B's vocabulary, landscape, reading list; never its conclusionsSession 6 · Tier 1 → Tier 2 · Unreliable, Dangerous at rung 4
7My field's tools — one bounded, checkable object. Never the finding.Outside Tier 1–3: a domain model with a domain verifier · Safe only where the verifier is external and you ran it · name the verifier before you run the tool, answer the five questions, record model version and date · governed by the Field-Frontier Tracker

Three rules: if you cannot name the verifier, the tool is not ready for your project · a benchmark number is not a deployment number — 87.3% became 0.33 on real patients [16] [17] · re-date the row every six months; inputs drift [37].

AI for ResearchersLandscape as of August 202627 / 33

Demo Preview — You Choose Two, Live

Five Field Demos Are Loaded. We Run the Two You Vote For.


AMedicine — PubMed-wide entity-and-relation search vs a chatbotwhich may well be right, and still cannot show you where it got it — free, no login [18]
BBiotech — a public sequence into a structure-prediction serverread the confidence colouring, find the disordered region — "it looks clean" is not evidence [19] [20]
CEngineering — text-to-CAD, a first-year part then a real oneand the vendor's own words for what happens next [40]
DPhysical sciences & maths — a false theorem, proved fluently, pasted into Leanthe kernel rejects it however good the prose; then a true one, and it accepts — the only demo in this series with a verdict [50]
ESocial sciences & humanities — LLM coding vs human coding, Cohen's κ liveten synthetic responses, a four-code book — and what a high κ still does not license [54] [56]

~Five minutes each; all five written out in the demo script with prompts, expected outcomes and fallbacks — the three we do not run are still yours. All sample data is synthetic.

AI for ResearchersLandscape as of August 202628 / 33

Closing the Series

Seven Sessions, Seven Artefacts, One Habit


"The proliferation of AI tools in science risks introducing a phase of scientific enquiry in which we produce more but understand less."

Messeri & Crockett, Nature 2024 — on illusions of understanding, and the scientific monocultures they hide [4]

seven pagesevery series artefact was a defence against that sentence — none of them names a product; today's is the Field-Frontier Tracker
the habitwhat would show me this is wrong, and did I go and look? — not new; just what research already was
and noticeAI's benefits in science are not evenly distributed [5] — whoever builds the tracker decides who keeps up. Make it a shared document.

This week: fill row 7 of your canvas for one real project, then build a Field-Frontier Tracker for your own field — and put a date on it.

AI for ResearchersLandscape as of August 202629 / 33

References · Cross-cutting · Medicine · Biotech

References (1–23)


  1. The Royal Swedish Academy of Sciences (2024). The Nobel Prize in Chemistry 2024.nobelprize.org/prizes/chemistry/2024/summary — accessed 2026-07-29
  2. The Royal Swedish Academy of Sciences (2024). The Nobel Prize in Physics 2024.nobelprize.org/prizes/physics/2024/summary — accessed 2026-07-29
  3. Wang, H., et al. (2023). Scientific discovery in the age of artificial intelligence. Nature 620, 47–60.doi.org/10.1038/s41586-023-06221-2 — accessed 2026-07-29
  4. Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature 627, 49–58.doi.org/10.1038/s41586-024-07146-0 — accessed 2026-07-29
  5. Gao, J., & Wang, D. (2024). Quantifying the use and potential benefits of AI in scientific research. Nature Human Behaviour 8, 2281–2292.doi.org/10.1038/s41562-024-02020-5 — accessed 2026-07-29
  6. The Royal Society (2024). Science in the age of AI.royalsociety.org/news-resources/projects/science-in-the-age-of-ai — accessed 2026-07-29
  7. U.S. Food and Drug Administration (2026). Artificial Intelligence-Enabled Medical Devices (list + downloadable file). Content current 2026-06-16.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices — accessed 2026-09-02
  8. U.S. Food and Drug Administration (2018). De Novo Decision Summary, DEN180001 (IDx-DR).accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf — accessed 2026-07-29
  9. Abràmoff, M. D., et al. (2018). Pivotal trial of an autonomous AI-based diagnostic system… npj Digital Medicine 1, 39.doi.org/10.1038/s41746-018-0040-6 — accessed 2026-07-29
  10. Wang, Z., et al. (2025). Systematic review and meta-analysis of regulator-approved deep learning systems for fundus DR detection. npj Digital Medicine.doi.org/10.1038/s41746-025-02223-8 — accessed 2026-07-29
  11. Antonissen, N., et al. (2026). AI in radiology: 173 commercially available products and their scientific evidence. European Radiology.doi.org/10.1007/s00330-025-11830-8 — accessed 2026-07-29
  12. Lawrence, R., et al. (2025). AI for diagnostics in radiology practice: a rapid systematic scoping review. eClinicalMedicine 83, 103228.pubmed.ncbi.nlm.nih.gov/40474995 — accessed 2026-07-29
  13. Goh, E., et al. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open 7(10), e2440969.doi.org/10.1001/jamanetworkopen.2024.40969 — accessed 2026-07-29
  14. Goh, E., et al. (2025). GPT-4 assistance for improvement of physician performance on patient care tasks: an RCT. Nature Medicine.doi.org/10.1038/s41591-024-03456-y — accessed 2026-07-29
  15. Qazi, I. A., Ali, A., Khawaja, A. U., et al. (2026). Automation Bias in LLM-Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — An RCT. NEJM AI 3(5).doi.org/10.1056/AIoa2501001 — accessed 2026-07-29
  16. Jin, Q., et al. (2024). Matching patients to clinical trials with large language models (TrialGPT). Nature Communications 15, 9074.doi.org/10.1038/s41467-024-53081-z — accessed 2026-07-29
  17. Kempf, E., et al. (2025). A prospective pragmatic evaluation of automatic trial matching tools in a molecular tumor board. npj Precision Oncology.doi.org/10.1038/s41698-025-00806-y — accessed 2026-07-29
  18. Wei, C.-H., et al. (2024). PubTator 3.0: an AI-powered literature resource… Nucleic Acids Research 52(W1), W540–W546.doi.org/10.1093/nar/gkae235 — accessed 2026-07-29
  19. Abramson, J., et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500.doi.org/10.1038/s41586-024-07487-w — accessed 2026-07-29
  20. Google DeepMind (2026). AlphaFold Server FAQ; alphafold3 repository; Model Parameters Terms of Use.alphafoldserver.com/faq (2026-07-29) · github.com/google-deepmind/alphafold3 — accessed 2026-09-02
  21. Yuan, R., et al. (2025). CASP16 Protein Monomer Structure Prediction Assessment. Proteins.doi.org/10.1002/prot.70031 — accessed 2026-07-29
  22. EMBL-EBI & Google DeepMind (2026). AlphaFold Protein Structure Database; with Varadi, M., et al. (2024). NAR 52, D368–D375.alphafold.ebi.ac.uk · doi.org/10.1093/nar/gkad1011 — accessed 2026-07-29
  23. Peng, C., et al. (2025). A comprehensive benchmarking of AlphaFold3… Briefings in Bioinformatics 26(6), bbaf616.doi.org/10.1093/bib/bbaf616 — accessed 2026-07-29
AI for ResearchersLandscape as of August 202630 / 33

References · Biotech · Engineering · Physical Sciences

References (24–45)


  1. Dauparas, J., et al. (2022). Robust deep learning–based protein sequence design using ProteinMPNN. Science 378, 49–56.doi.org/10.1126/science.add2187 — accessed 2026-07-29
  2. Watson, J. L., et al. (2023). De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100.doi.org/10.1038/s41586-023-06415-8 — accessed 2026-07-29
  3. Pacesa, M., et al. (2025). One-shot design of functional protein binders with BindCraft. Nature 646(8084), 483–492.doi.org/10.1038/s41586-025-09429-6 — accessed 2026-07-29
  4. Lauko, A., Ahern, W., et al. (2026). Atom-level enzyme active site scaffolding using RFdiffusion2. Nature Methods 23, 96–105.doi.org/10.1038/s41592-025-02975-x — accessed 2026-07-29
  5. Hayes, T., et al. (2025). Simulating 500 million years of evolution with a language model. Science 387, eads0018.doi.org/10.1126/science.ads0018 — accessed 2026-07-29
  6. Jayatunga, M. K. P., et al. (2024). How successful are AI-discovered drugs in clinical trials? Drug Discovery Today 29(6), 104009.doi.org/10.1016/j.drudis.2024.104009 — accessed 2026-07-29
  7. Insilico Medicine, et al. (2025). A generative AI-discovered TNIK inhibitor for IPF: a randomized phase 2a trial. Nature Medicine.doi.org/10.1038/s41591-025-03743-2 — accessed 2026-07-29
  8. Kedzierska, K. Z., et al. (2025). Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biology 26, 101.doi.org/10.1186/s13059-025-03574-x — accessed 2026-07-29
  9. McGreivy, N., & Hakim, A. (2024). Weak baselines and reporting biases lead to overoptimism in ML for fluid-related PDEs. Nature Machine Intelligence 6.doi.org/10.1038/s42256-024-00897-5 · arxiv.org/abs/2407.07218 — accessed 2026-07-29
  10. Krishnapriyan, A. S., et al. (2021). Characterizing possible failure modes in physics-informed neural networks. NeurIPS 2021.arxiv.org/abs/2109.01050 — accessed 2026-07-29
  11. Azizzadenesheli, K., et al. (2024). Neural operators for accelerating scientific simulations and design. Nature Reviews Physics 6, 320–328.doi.org/10.1038/s42254-024-00712-5 — accessed 2026-07-29
  12. Lam, R., et al. (2023). Learning skillful medium-range global weather forecasting (GraphCast). Science 382, eadi2336.doi.org/10.1126/science.adi2336 — accessed 2026-07-29
  13. ECMWF (2026). AIFS Machine Learning data; and Haiden & Chevallier (2026), Forecast performance 2025, Newsletter 187.ecmwf.int/en/forecasts/dataset/aifs-machine-learning-data — accessed 2026-07-29
  14. Ben Bouallègue, Z., Raoult, B., & Chantry, M. (2026). Farewell to the external AI models. ECMWF AIFS Blog, 11 May 2026.doi.org/10.21957/fa8ad01483 — accessed 2026-07-29
  15. Ohana, R., et al. (2024). The Well: a large-scale collection of diverse physics simulations for ML. NeurIPS 2024 D&B.arxiv.org/abs/2412.00568 · polymathic-ai.org/the_well — accessed 2026-07-29
  16. Autodesk (2020; product page 2026). Topology Optimization is not Generative Design; Generative design for manufacturing.autodesk.com/products/fusion-360/blog · autodesk.com/solutions/generative-design/manufacturing — accessed 2026-07-29
  17. Zoo (2026). Frequently Asked Questions; Text-to-CAD.zoo.dev/docs/faq · zoo.dev/design-studio (v1.4.4) — accessed 2026-09-02
  18. Pearce, H., et al. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. IEEE S&P 2022, 754–768.doi.org/10.1109/SP46214.2022.9833571 · arxiv.org/abs/2108.09293 — accessed 2026-07-29
  19. Becker, J., et al. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Preprint.arxiv.org/abs/2507.09089 — accessed 2026-07-29
  20. Merchant, A., et al. (2023). Scaling deep learning for materials discovery (GNoME). Nature 624, 80–90.doi.org/10.1038/s41586-023-06735-9 — accessed 2026-07-29
  21. Cheetham, A. K., & Seshadri, R. (2024). Artificial Intelligence Driving Materials Discovery? Chemistry of Materials 36(8), 3490–3495.doi.org/10.1021/acs.chemmater.4c00643 — accessed 2026-07-29
  22. Szymanski, N. J., et al. (2023). An autonomous laboratory for the accelerated synthesis of inorganic materials (A-Lab). Nature 624, 86–91. Read with [69].doi.org/10.1038/s41586-023-06734-w — accessed 2026-07-29
AI for ResearchersLandscape as of August 202631 / 33

References · Physical Sciences & Mathematics · Social Sciences

References (46–58)


  1. Leeman, J., et al. (2024). Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. PRX Energy 3, 011002.doi.org/10.1103/PRXEnergy.3.011002 — accessed 2026-07-29
  2. Riebesell, J., et al. (2025). A framework to evaluate machine learning crystal stability predictions (Matbench Discovery). Nature Machine Intelligence 7, 836–847.doi.org/10.1038/s42256-025-01055-1 · matbench-discovery.materialsproject.org — accessed 2026-09-02
  3. NSF–DOE Vera C. Rubin Observatory (2026). Rubin Observatory Launches Real-Time Alerts (25 Feb 2026); Alerts and brokers.rubinobservatory.org/news/first-alerts — accessed 2026-07-29
  4. Hubert, T., Masoom, R., Barekatain, M., et al. / Google DeepMind (2025). Olympiad-level formal mathematical reasoning with reinforcement learning (AlphaProof). Nature.doi.org/10.1038/s41586-025-09833-y — accessed 2026-07-29
  5. Lean community (2026). Lean 4 Web; LeanSearch; Mathlib statistics.live.lean-lang.org · leansearch.net · leanprover-community.github.io/mathlib_stats.html — accessed 2026-09-02
  6. Epoch AI (2026). FrontierMath; FrontierMath Tiers 1–4 (v2, 12 June 2026).epoch.ai/frontiermath — accessed 2026-07-29
  7. IMO 2025 Organisers / Dolinar, G. (2025). The 66th International Mathematical Olympiad draws to a close today, 19 July 2025.imo2025.au/news/the-66th-international-mathematical-olympiad-draws-to-a-close-today — accessed 2026-07-29
  8. Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. PNAS 120(30), e2305016120.doi.org/10.1073/pnas.2305016120 — accessed 2026-07-29
  9. Egami, N., Hinck, M., Stewart, B. M., & Wei, H. (2023). Using Imperfect Surrogates for Downstream Inference (DSL). NeurIPS 2023.arxiv.org/abs/2306.04746 · naokiegami.com/dsl — accessed 2026-07-29
  10. Ollion, É., Shen, R., Macanovic, A., & Chatelain, A. (2024). The dangers of using proprietary LLMs for research. Nature Machine Intelligence 6, 4–5.doi.org/10.1038/s42256-023-00783-6 — accessed 2026-07-29
  11. Liu, X., et al. (2025). Qualitative Coding with GPT-4: Where it Works Better. J. Learning Analytics 12(1), 169–185.doi.org/10.18608/jla.2025.8575 — accessed 2026-07-29
  12. Argyle, L. P., et al. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31(3), 337–351.doi.org/10.1017/pan.2023.2 — accessed 2026-07-29
  13. Bisbee, J., et al. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis 32(4), 401–416.doi.org/10.1017/pan.2024.5 — accessed 2026-07-29
AI for ResearchersLandscape as of August 202632 / 33

References · Social Sciences & Humanities · Frontier Venues

References (59–70)


  1. Boelaert, J., et al. (2025). Machine Bias: How Do Generative Language Models Answer Opinion Polls? Sociological Methods & Research.doi.org/10.1177/00491241251330582 — accessed 2026-07-29
  2. Park, J. S., et al. (2026). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals. Preprint, v3, never journal-published.arxiv.org/abs/2411.10109 — accessed 2026-09-02
  3. Vesuvius Challenge (2026). An entire Herculaneum scroll has been read for the first time, 25 June 2026.scrollprize.org/firstscroll · scrollprize.org/data — accessed 2026-07-29
  4. READ-COOP SCE / Transkribus (2026). Pricing; Character Error Rate (CER) explained.transkribus.org/pricing · transkribus.org/character-error-rate-cer-explained — accessed 2026-07-29
  5. Hutchinson, D. (2026). Benchmarking as Source Criticism: From Recognition to Reasoning in LLM Assessment. J. Open Humanities Data 12, art. 43.doi.org/10.5334/johd.489 — accessed 2026-07-29
  6. MAXQDA/VERBI, Lumivero (NVivo), ATLAS.ti, Taguette (2026). AI feature documentation.maxqda.com/products/ai-assist · lumivero.com/products/nvivo · atlasti.com · taguette.org/about.html — accessed 2026-07-29
  7. NEJM Group (2026). About NEJM AI. ISSN 2836-9386.ai.nejm.org/about — accessed 2026-07-29
  8. Protein Structure Prediction Center (2026). CASP16.predictioncenter.org/casp16 — accessed 2026-07-29
  9. Summer Institutes in Computational Social Science (2026). About SICSS.sicss.io/about — accessed 2026-07-29
  10. Jacobson, R. D. (2025). The AI drug revolution needs a revolution. npj Drug Discovery.doi.org/10.1038/s44386-025-00013-6 — accessed 2026-07-29
  11. Szymanski, N. J., et al. (2026). Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature 650, E1.doi.org/10.1038/s41586-025-09992-y — accessed 2026-07-29
  12. Kim, D., Woodbury, S. M., Ahern, W., et al. (2025). Computational design of metallohydrolases. Nature 649, 246–253.doi.org/10.1038/s41586-025-09746-w — accessed 2026-08-08