← series index
# Session 7: Discipline Deep-Dives: Specialized AI by Field
*AI for Researchers — attendee handout · Landscape as of August 2026*
---
## At a glance
| | |
|---|---|
| **Session** | 7 of 7 — Discipline Deep-Dives: Specialized AI by Field |
| **You'll learn** | What the non-chatbot AI reshaping five disciplines actually does, what each landmark system's own authors say it does **not** do, and how to keep tracking your field after this series ends. |
| **You'll practise** | Two of five field demos, chosen live by poll — all five are written out in `demo-script.md`. |
| **Prerequisite** | None. Uses Session 1's zones and tiers; links Session 4's policy table; **completes row 7** of Session 5's canvas. |
| **Tools referenced** | None endorsed. Every tool named here is named for a capability, with its own documented limit beside it. |
---
## The one-paragraph version
Specialized AI is not a better chatbot but a different object, and it was decisive before the chatbots arrived — the 2024 Nobel Prizes in Physics and Chemistry both went to it [1][2]. It outputs a **domain object** (a coordinate file, a candidate crystal, a proof term, a coded transcript) rather than prose, so you cannot judge it by reading it. It is usually scored by somebody else — CASP [66], Matbench Discovery [47], a regulator's premarket review [7], an olympiad paper [49] — and it hands off to a verifier that is not software: a wet lab, a telescope, a proof kernel, a second human coder. **That verifier is the entire safety story.** Where it is cheap and external, the field is healthy: fold prediction is "nearly solved" [21], and a Herculaneum scroll nobody could open has been read end to end [61]. Where it is expensive, slow, or run by the claimants, corrections take years and arrive in print — two peer-reviewed rebuttals on one celebrated materials result [44][46], then a journal Author Correction restating it as "36 compounds from a set of 57 targets" [69]. Between benchmark and deployment the numbers move hard — trial matching scored 87.3% on synthetic patients [16] and precision 0.33 on 157 real ones [17]. The transferable skill is not a tool list. It is five questions, and the habit of asking who is allowed to check.
---
## The Five Questions — ask these of any field-specific AI tool
Write these on the inside cover of your notebook. The tool names will change by Christmas; these will not.
1. **What exactly does it output, and what does its own paper say it does *not* do?** Every honest landmark system publishes a limitations section. AlphaFold 3's names a **4.4% chirality violation rate** and hallucinated order in disordered regions [19].
2. **Who scored it, and were they independent of the people who built it?** A vendor benchmark and a blinded community assessment are not the same evidence [66][47].
3. **What is the gap between the benchmark and the deployment?** Same task, two settings: **87.3% → 0.33** [16][17].
4. **What is the external verifier, and can you afford to run it?** Ninety-six wells, a telescope night, a proof kernel, a second coder. If you cannot run it, the output is not a finding.
5. **What happens when the inputs shift?** When ECMWF upgraded its physics model, the fine-tuned ML forecasters **lost skill** while the physics gained it [37].
---
## AI frontier maps — one per field
Each map: what to actually use · one seminal thing to read · where to follow updates · the caution.
### Medicine & clinical research
| | |
|---|---|
| **Named tools** | **PubTator 3.0** (NCBI) — entity and relation search over ~36M abstracts and 6M full texts; free, no login [18]. **AlphaFold DB** for target work [22]. **Clinical-trial matching services** — useful for *ranking and explaining*, never for determining eligibility; the 87.3%/0.33 evidence below is from **2024–25 systems**, and we found no comparable real-patient evaluation of a 2026-generation matcher [16][17]. |
| **Read this one** | Antonissen et al. (2026), *AI in radiology: 173 commercially available products and their scientific evidence*, *European Radiology* [11]. The most honest single picture of the evidence base: peer-reviewed evidence rose 36% → 66%, but only **31%** reaches clinical decisions, outcomes or cost, and prospective designs went **19% → 16%**. |
| **Follow updates** | FDA's **AI-Enabled Medical Device List**, updated in place [7] · **NEJM AI**, which pairs pre-clinical with clinical work [65] · your specialty registry. |
| **Caution** | **Automation bias survives AI training.** A 2025 trial on a GPT-4-class assistant: 44 physicians who had completed a 20-hour AI-literacy course, 264 cases, and exposure to erroneous LLM suggestions cut composite accuracy from **84.9% to 73.3%**, adjusted **−14.0 points** (95% CI −18.9 to −9.1, *P* < 0.0001) [15]. In the 2025 GPT-4 RCT the model alone was statistically indistinguishable from the model-plus-physician (−0.9%, *P* = 0.8) [14]; in the 2024 RCT the model alone beat only the conventional-resources arm it was tested against [13]. All three trials ran on the **2023–25 model cohort** and none has been repeated on a 2026-generation assistant — which is why the effect to plan for is *automation bias*, a property of how people use the output, not a property of any model version. |
| **Regulatory note** | Device authorisation [7] governs the **product**; **Session 4's Publisher AI-Policy Comparison Table** governs **your manuscript** — different regimes, and you are subject to both. Check the Session 4 table before writing up anything built with these tools. |
### Biotechnology & life sciences
| | |
|---|---|
| **Named tools** | **AlphaFold Server** — proteins, nucleic acids, ligands, ions, modifications; 30 jobs/day, closed ligand list, a 2021-09-30 template cutoff that is only a changeable **default** (quotas as documented 2026-07-29), **weights gated and non-commercial** — those weights terms re-read 2026-09-02 and unchanged [20]. **AlphaFold DB** — >200M predictions, free, no login [22]. **ProteinMPNN** (sequence for a given backbone, 52.4% recovery vs Rosetta's 32.9% [24]) and **RFdiffusion** (the backbone itself [25]) — different jobs. |
| **Read this one** | Yuan et al. (2025), *CASP16 Protein Monomer Structure Prediction Assessment*, *Proteins* [21]. Independent, blinded, and it tells you both halves: fold prediction is "nearly solved", and "**model ranking remains a persistent weakness across most groups**". |
| **Follow updates** | **CASP** rounds and their assessment papers [66][21] · AlphaFold DB release notes [22] · bioRxiv, then the journal version. |
| **Caution** | **The wet lab still decides, and the spread is the headline.** De novo binder success: **19%** across five targets, 95 designs tested per target [25]; **10–100%, averaging 46.3%**, across twelve targets in another study [26] — whose "<0.1%" Rosetta figure describes prior literature, not a matched comparison [26]. RFdiffusion2 scaffolds all 41 benchmark active sites against 16 for RFdiffusion, but that is an **in-silico** result; experimentally it found actives after testing "fewer than 96 sequences" [27] — its best design, second round, with an explicit general base, reached **53,000 ± 5,000 M⁻¹s⁻¹** [70]. On drug discovery, read the caveat inside the most positive clinical paper of 2025: "AI-discovered drugs have experienced **similar levels of phase 2 trial failure** as non-AI-discovered drugs, and **none has so far progressed through phase 3**" [30] — still true: the first registered phase 3 (NCT07687459, rentosertib, 320 patients) was still "not yet recruiting" on 2026-09-02, record last updated 1 July 2026. The often-quoted 80–90% Phase I rate rests on **24 molecules, 21 successes** [29], challenged in print [68]. |
### Engineering
| | |
|---|---|
| **Named tools** | **Neural surrogates / neural operators** for PDEs — real, best evidenced in weather: ECMWF's **AIFS** operational since 25 Feb 2025 (ensemble from 1 July 2025), both to v2 on 12 May 2026 [36]. **GraphCast** — 10-day forecasts at 0.25° in under a minute, beating the best operational system on 90% of 1,380 targets, but does **not** assimilate observations [35]. **Generative design** in commercial CAD and **text-to-CAD** — two different things (below). **The Well** — 15 TB, 16 datasets, for testing against a strong baseline [38]. |
| **Read this one** | McGreivy & Hakim (2024), *Weak baselines and reporting biases lead to overoptimism in ML for fluid-related PDEs*, *Nature Machine Intelligence* [32]. **79% (60 of 76)** of papers claiming to beat a numerical method compared against a weak baseline. Read it before believing any speedup number — including "four to five orders of magnitude", which comes from **the developers of neural operators reviewing their own method** [34], on timings of exactly the kind [32] finds weakly baselined. |
| **Follow updates** | Open benchmark suites such as The Well [38] · **ECMWF's AIFS pages and blog**, which publish their failures as well as their wins [36][37]. |
| **Caution** | **"Generative design" has no text prompt in it.** The incumbent vendor's inputs are **preserve geometry** ("a body you want to include in the final shape") and **obstacle geometry** ("a body you want to exclude"), plus loads, constraints, materials and manufacturing methods [39] — verified July 2026; the vendor blocked a 2026-09-02 re-check and has announced prompt-driven CAD as forthcoming, so **re-date this claim before you repeat it**. Text-to-CAD is separate and younger; at **v1.4.4 (28 Aug 2026)** its vendor still calls the model "still experimental" and warns its agent may "produce incorrect geometry, or suggest designs that aren't manufacturable or safe", while the documented output has shifted to editable **KCL** project files with **STEP** export a separate step [40]. On code, two **dated** baselines: on "Early-2025 AI", 16 experienced developers measured **19% slower** while believing they were 20% faster [42, preprint]; on **2022-era GitHub Copilot**, ~**40%** of generated programs in 89 security scenarios were vulnerable [41]. Neither has been re-run on the 2026 generation — quote them with the year attached. |
### Physical sciences & mathematics
| | |
|---|---|
| **Named tools** | **Matbench Discovery** — a live, open leaderboard on a blind test set; the peer-reviewed paper says models can "effectively and cheaply pre-screen" candidates, and the project's README puts it more crisply as "robust enough to deploy them as triaging steps" — **triage, not discovery**. Snapshot 2026-09-02: 42 models, led by EquiformerV3+DeNS-OAM (F1 0.931), added 7 April 2026 — *and that sentence is stale by the time you read it, which is the point of using a leaderboard rather than a headline* [47]. **Alert brokers** — seven full-stream plus two down-stream services on Rubin Observatory's stream, using ML to filter, sort and classify up to **seven million alerts per night** [48]. **Lean 4 + Mathlib** — 286,514 theorems from 772 contributors on 2026-09-02, every one machine-checked, and the counts advance daily [50]; **LeanSearch** to find a lemma [50]. |
| **Read this one** | Read the pair in order: Merchant et al. (2023) [43], then **Cheetham & Seshadri (2024)**, *Chemistry of Materials* [44]. Thirty minutes, and you will never again read a discovery claim without asking who re-examined it. Then Leeman et al. (2024), *PRX Energy* [46], and the journal **Author Correction** [69] restating the autonomous-lab result as "36 compounds from a set of 57 targets" and conceding the materials were "new to the prediction platform, not necessarily new to science". |
| **Follow updates** | Matbench Discovery leaderboard [47] · Rubin broker documentation [48] · Mathlib and the Lean Zulip [50] · and **benchmark changelogs** — in June 2026 one research-mathematics benchmark shipped a v2 "addressing errors in **42% of problems**" [51]. |
| **Caution** | **A predicted compound is not a material,** and the rebuttals say so in print: "scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility" [44][46]. In mathematics, distinguish a **machine-checked** proof from a persuasive one: AlphaProof solved three of five non-geometry IMO 2024 problems from statements **manually formalised by experts**, each needing 2–3 days of compute, both combinatorics problems unsolved [49] — while the IMO says of 2025 claims it "**cannot validate the methods**… or whether the results can be reproduced" [52]. The 2026 olympiad (Shanghai, 10–21 July 2026) has been held since; **as of the 2025 cycle that statement is still the IMO's own last word**, so check for a newer one before quoting it. |
### Social sciences & humanities
| | |
|---|---|
| **Named tools** | **LLM annotation against a fixed codebook** — in 2023, ChatGPT was cheap and better than **MTurk crowd work**, +~25 points at under **$0.003** per annotation — though absolute accuracy was a modest **59–83%**, and "beats trained annotators" is about intercoder *agreement*, not accuracy [53]. **QDA software AI features** (MAXQDA, NVivo, ATLAS.ti) — coding *suggestions* you accept or reject; none claims validated codes or publishes an independent inter-rater evaluation [64]. **Taguette** — free, open source, **no** AI features: the deliberate choice for sensitive or IP-restricted material [64]. **Transkribus** for handwritten text recognition — free tier 50 credits/month [62]. |
| **Read this one** | Egami, Hinck, Stewart & Wei (2023), *Using Imperfect Surrogates for Downstream Inference* [54]. It is the paper that will save you a retraction: direct use of surrogate labels in downstream analyses gives "**substantial bias and invalid confidence intervals, even with high surrogate accuracy of 80–90%**", because model errors correlate with your covariates. |
| **Follow updates** | *Sociological Methods & Research* and *Political Analysis* for the method debates [58][59] · *Journal of Open Humanities Data* [63] · **SICSS**, free computational-social-science training [67]. |
| **Caution** | **Manifest codes transfer; interpretive ones do not.** Measured on **GPT-4**: κ ≥ 0.70 for 25 of 34 constructs — only when the best prompting strategy was chosen *per construct*, zero-shot ranging from **κ = 0.91 to κ = 0.11**; GPT-4 struggled most with the constructs human coders struggled with too, human–human κ itself running **0.24–0.87** [56]. Expect the current models to score better; expect the *shape* to hold, because **a better model raises the κ and does not make an ambiguous construct unambiguous** — and note that Egami's bias result above is triggered by *high* accuracy, not excused by it [54]. On synthetic respondents, read the whole exchange as the record of a debate fought on **2022–24 models** — GPT-3 for the claim [57], GPT-3.5 for the first rebuttal [58] — which nobody has re-run on the 2026 generation. What the rebuttals found is structural: they match the mean and lose the variance — "less variation in responses than in the real surveys, and regression coefficients often differ significantly" [58], replicated as "strong bias and a low variance on each topic" varying "randomly… from one topic to the next" [59]. The best-known simulation paper remains an **unpublished preprint** (v3, revised 28 June 2026, checked 2026-09-02); a structured survey scores within one point of a two-hour interview [60]. Record your model version and date: repeating an experiment weeks later "often yielded different results" [55]. |
---
## The Field-Frontier Tracker — fillable
One page, **one field**, one date. Fill columns 3 and 4 first: if you cannot name the verifier or the place where results get contested in public, you have a tool, not a frontier.
*The maps above were last re-dated **2026-09-02**; the fast-moving rows (leaderboard snapshots, tool versions, regulator counts, trial statuses) carry their own check dates inline. Rule (3) below applies to this handout too.*
**Field:** ______________________ **Date filled:** __________ **Re-check on:** __________
| | Your answer |
|---|---|
| **Landmark system** — what it outputs | |
| **Its own stated limit** (quote its paper or docs) | |
| **Who scored it** — and were they independent? | |
| **The external verifier** — and can you afford it? | |
| **Where the frontier is published** — venue, benchmark, leaderboard, changelog | |
| **One review article to hand a new PhD student** | |
---
## Row 7 of your Personal AI Workflow Canvas — the last row
| Lifecycle stage | What I hand to AI | Tier · Zone | My non-negotiable check | Artefact |
|---|---|---|---|---|
| **7 · My field's tools** (S7) | One **bounded, checkable object**: a predicted structure, a candidate list, a coded transcript, a formal proof obligation. **Never the finding.** | Outside Tier 1–3: a domain model with a domain verifier · ☐ Safe ☐ Unreliable ☐ Dangerous | ☐ verifier named **before** running the tool ☐ five questions answered ☐ benchmark-vs-deployment gap checked ☐ model version + date recorded | **The Field-Frontier Tracker** |
**My verifier is:** ______________________ **It costs:** ______________________ **I can/cannot run it:** __________
**Three rules for row 7.** (1) If you cannot name the verifier, the tool is ready for your reading list, not your project. (2) A benchmark number is not a deployment number [16][17]. (3) Re-date the row every six months — inputs drift, and fine-tuned models lose skill when they do [37].
---
## Series-wide closing checklist
Seven sessions, seven artefacts, one page. Before AI output enters any part of your next project:
- [ ] **S1 — Capability/Failure Map.** I have placed this task in **Safe / Unreliable / Dangerous**, and I remember Rule 2: rarity moves tasks rightward.
- [ ] **S1 — Tool Landscape Taxonomy.** I know whether I am using **Tier 1** (general chatbot), **Tier 2** (research-specific), **Tier 3** (deep-research agent) — or a specialised tool that sits outside all three.
- [ ] **S1 — CRIT.** My prompt has **C**ontext, **R**ole and register, **I**nstructions and constraints, **T**ask — and I iterated.
- [ ] **S2 — Literature-Discovery Tool Matrix.** Every citation resolves in a database before I use it. No exceptions.
- [ ] **S3 — Summary-Trust Triage.** I know which summaries oblige me to reopen the paper, and I reopened them.
- [ ] **S4 — Publisher AI-Policy Comparison Table.** I have checked my **target venue's** policy and written the disclosure it asks for.
- [ ] **S5 — Personal AI Workflow Canvas + Four Reproducibility Guardrails.** Code kept, versions pinned, seed set, one number recomputed by hand.
- [ ] **S6 — Cross-Field Orientation Ladder.** If I crossed a field boundary: glossary checked against one field-B review, every paper resolved, two of eight read in full, one field-B researcher read the memo.
- [ ] **S7 — Field-Frontier Tracker.** I named the external verifier **before** I ran the tool, and I ran it.
- [ ] **The one habit under all seven:** I can say what would show me this is wrong — and I went and looked.
---
## Further reading
- Wang, H., et al. (2023). *Scientific discovery in the age of artificial intelligence*. Nature 620, 47–60. https://doi.org/10.1038/s41586-023-06221-2 — the standing cross-disciplinary review; the map this session tours. [3]
- Messeri, L., & Crockett, M. J. (2024). *Artificial intelligence and illusions of understanding in scientific research*. Nature 627, 49–58. https://doi.org/10.1038/s41586-024-07146-0 — "produce more but understand less"; the closing argument of the series. [4]
- Cheetham, A. K., & Seshadri, R. (2024). *Artificial Intelligence Driving Materials Discovery?* Chem. Mater. 36(8), 3490–3495. https://doi.org/10.1021/acs.chemmater.4c00643 — the best-written rebuttal in this deck. A model for checking a discovery claim. [44]
- Egami, N., et al. (2023). *Using Imperfect Surrogates for Downstream Inference*. NeurIPS 2023. https://arxiv.org/abs/2306.04746 — read even if you never code a transcript; the bias argument generalises to any AI-generated variable. [54]
- McGreivy, N., & Hakim, A. (2024). *Weak baselines and reporting biases…* Nature Machine Intelligence 6, 1256–1269. https://doi.org/10.1038/s42256-024-00897-5 — how a literature became overoptimistic without anyone lying. [32]
- The Royal Society (2024). *Science in the age of AI*. https://royalsociety.org/news-resources/projects/science-in-the-age-of-ai/ — first-party, drawn on 100+ scientists; four recommendations worth quoting at your department. [6]
- Gao, J., & Wang, D. (2024). *Quantifying the use and potential benefits of AI in scientific research*. Nature Human Behaviour 8, 2281–2292. https://doi.org/10.1038/s41562-024-02020-5 — AI's benefits in science are not evenly distributed; whoever builds the tracker decides who keeps up. [5]
*Full annotated source list — 70 sources, with the verified wording and the caveat behind every number: see `sources.md`. All five field demos, including the three not run live: see `demo-script.md`.*
---
*AI for Researchers · Session 7: Discipline Deep-Dives: Specialized AI by Field · Landscape as of August 2026*