AI for Researchers

Session 7: Discipline Deep-Dives: Specialized AI by Field

Five fields, five landmark systems, and the one question that tells you whether any of them is safe to use

{{Presenter Name}}

Landscape as of August 2026

2-Minute Recap · The Whole Series

The Series So Far


Today (Session 7): the AI that is not a chatbot — the systems reshaping five disciplines, what each one really does, and how to track yours after this series ends.

Learning Objectives

By the End of This Session You Will Be Able To…


  1. Identify the specialized, non-chatbot AI tools transforming your own discipline.
  2. Explain at least one specialized AI application in each of medicine, life sciences, engineering, physical sciences/math, and social sciences/humanities.
  3. Evaluate a specialized, field-specific AI tool relevant to your own research area.
  4. Apply strategies to track your field's AI frontier — key venues, benchmarks, and review articles.
  5. Compare specialized discipline-specific tools to the general-purpose tools covered earlier in the series.

Section 01

The Tier the Taxonomy Left Out

Six sessions on tools that write. This one is about tools that predict a structure, propose a compound, or refuse a proof — and why they fail in a completely different direction.

Literacy Foundation — The Key Slide

Specialized AI Is Not a Better Chatbot. It Is a Different Object.


So Session 1's Tier 1 / 2 / 3 sorted tools by how general they are. Today's question is different and better: what checks it, and did you run the check?

Series Artefact — The Reusable Part of Today

Five Questions to Ask Any Field-Specific AI Tool


  1. What exactly does it output, and what does its own paper say it does not do? Every honest landmark system publishes a limitations section. AlphaFold 3's names a 4.4% chirality violation rate and hallucinated order in disordered regions. [19]
  2. Who scored it, and were they independent of the people who built it? A vendor benchmark and a blinded community assessment are not the same evidence. [66] [47]
  3. What is the gap between the benchmark and the deployment? Trial-matching scored 87.3% on synthetic patients [16] and precision 0.33 on 157 real ones. [17] Same task.
  4. What is the external verifier, and can you afford to run it? Ninety-six wells, a telescope night, a proof kernel, a second coder. If you cannot run it, you cannot use the output as a finding.
  5. What happens when the inputs shift? When ECMWF upgraded its physics model, the fine-tuned ML forecasters lost skill while the physics gained it. [37]

These five travel across all five fields, and they are the spine of the tracker on slide 26. Write them on the inside cover of your notebook; the tool names will change by Christmas.

Section 02

Medicine & Clinical Research

The only field on today's tour where a regulator already decides what the software is allowed to claim — and where the published numbers fall furthest between the benchmark and the clinic.

Medicine — Landmark Systems, Regulatory Context

1,524 Authorised Devices, and Almost None of Them Is What You Picture


1,524entries counted by us in the 2026-06-16 download of FDA's AI-Enabled Medical Device List; decisions spanning 1995 to 30 March 2026 [7]
76.4%of those entries have Radiology as lead review panel — our count, same download. FDA publishes none of these numbers as prose [7]
0identified as large-language-model based — and still 0: FDA's sentence is future tense, it "will explore methods to identify and tag medical devices that incorporate foundation models" [7]

Medicine — What Changed for Practitioners

The Number That Travels, and the Number That Does Not


87.3% → 0.33Clinical-trial matching, 2024–25 systems. TrialGPT: criterion-matching accuracy 87.3% on 1,015 pairs, screening time down 42.6% — on 183 synthetic patients [16]. Four deployed tools on 157 real tumour-board patients: mean precision 0.33, recall 0.32, and 38% of patients got no trial at all [17]. No comparable real-patient evaluation of a 2026-generation matcher was found — the gap, not the number, is what transfers
36% → 66%, but 31%Imaging. Across 173 CE-certified radiology AI products, the share with any peer-reviewed evidence rose from 36% to 66% — yet only 31% have evidence at the level of clinical decisions, outcomes or cost, and prospective designs went 19% → 16% [11]

Medicine — The Caution

Twenty Hours of AI Training Did Not Protect Them


"Physicians demonstrate substantial automation bias when exposed to erroneous LLM recommendations, even with voluntary consultation and prior AI literacy training." Randomised trial run in 2025 on a GPT-4-class assistant, 44 physicians who had completed a 20-hour AI-literacy course, 264 cases: composite diagnostic accuracy fell from 84.9% to 73.3%, adjusted −14.0 points (95% CI −18.9 to −9.1, P < 0.0001) [15]

Section 03

Biotechnology & Life Sciences

The field with the clearest win of the decade, the most disciplined published limitations, and the most contested commercial promise. All three are on the next three slides.

Biotech — Landmark System, From Its Own Paper

AlphaFold 3: What It Does, and What Its Authors Say It Does Not


Does

  • Predicts the joint structure of complexes containing proteins, nucleic acids, small molecules, ions and modified residues — one model, not protein-only [19]
  • Beats classical docking without structural inputs (vs Vina, P = 2.27 × 10⁻¹³) [19]
  • Underwrites a database of over 200 million predicted structures, free and bulk-downloadable [22]
  • Blinded verdict, CASP16: single-domain fold prediction is "nearly solved" [21]

Does not

  • Give dynamics: static outputs "as seen in the PDB"; multiple seeds "do not approximate the solution ensemble" [19]
  • Guarantee chemistry: 4.4% chirality violation rate; hallucinated order in disordered regions, without AF2's visual tell [19]
  • Rank outputs: "model ranking remains a persistent weakness across most groups" [21]
  • Come without strings: 30 jobs/day, closed ligand list, 2021-09-30 cutoff (changeable default); weights gated, non-commercial [20]

Every line on both sides comes from the AF3 paper [19], DeepMind's own docs [20], or the independent CASP16 assessment [21] [23] — never from coverage. Quotas as documented 2026-07-29; the gated, non-commercial weights terms re-read 2026-09-02, unchanged. An AF3 model that looks clean is not evidence that it is.

Biotech — What Changed for Practitioners

Design Moved From Reading Proteins to Writing Them — at a Measured Cost


The design software is free; the verifier is your consumables budget — question 4 on slide 6. "41 of 41" is computational; "19%" and "46.3%" are experimental. Never compare across that line.

Biotech — Where the Story Is Contested

AI Drug Discovery: State the Claim, Then State Its Own Authors' Caveat


"In Phase I we find AI-discovered molecules have an 80–90% success rate, substantially higher than historic industry averages… In Phase II the success rate is ~40%, albeit on a limited sample size, comparable to historic industry averages." Jayatunga, Ayers, Bruens, Jayanth & Meier, Drug Discovery Today 2024 — on 24 molecules, 21 successes. The authors call this "early signs of potential", not proof [29]

Section 04

Engineering

One operational triumph, one measured literature-wide bias, and a word — "generative design" — that does not mean what the audience assumes it means.

Engineering — Landmark System, and the Published Correction

Simulation Surrogates: Real, Operational, and Systematically Oversold


25 Feb 2025ECMWF's ML forecast system AIFS Single goes operational; the ensemble follows 1 July 2025; both to v2 on 12 May 2026 [36]
5–15%the honest accuracy gain: medium-range error reduction vs the physics-based IFS. The order-of-magnitude win is in cost, not skill [36]
79%of ML-for-fluid-PDE papers claiming to beat a numerical method compared against a weak baseline (60 of 76) [32]

Engineering — What Changed, and the Caution

"Generative Design" Is Not a Text Prompt — and Your Model Ages


[42] is still a preprint (carried from Session 5); neither code number has been re-measured on the 2026 models. Dataset shift is not hypothetical: it is why a surrogate validated this year needs revalidating when its upstream data changes — and why open benchmarks exist. [38]

Section 05

Physical Sciences & Mathematics

Where two disciplines built the best answer anyone has to AI hype: they let outsiders re-examine the claim. One of them got a correction. The other could not even try.

Physical Sciences — Where the Story Is Contested

Materials Discovery: the Claim, and the Peer-Reviewed Rebuttal


2.2M · 380kThe claim (Nature, 2023). GNoME predicted 2.2 million structures, ~380,000 stable — "an order-of-magnitude expansion in stable materials known to humanity" [43]. A companion lab reported synthesising novel materials over 17 days [45]
0 · 0The rebuttals (peer-reviewed, 2024). "Scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility… we have yet to find any strikingly novel compounds" [44]. On the autonomous lab, re-examining all 43 products: "no new materials have been discovered in that work" [46]

The live answer is a leaderboard, not a headline: Matbench Discovery's peer-reviewed hedge is that models "effectively and cheaply pre-screen" candidates; the crisper "triaging steps" line is the project's GitHub README, not the journal. [47] Snapshot, 2026-09-02: 42 models, led by EquiformerV3+DeNS-OAM (F1 0.931), added April 2026 — stale the moment it is read.

Physical Sciences — What Changed for Practitioners

In Observational Physics, AI Is Not the Discovery. It Is the Only Way to Get to It.


"Rubin issued 800,000 alerts the night of 24 February… a system expected to eventually produce up to seven million alerts per night… scientists rely on a network of intelligent software platforms known as brokers. These systems use machine learning algorithms to filter, sort, and classify the alerts." NSF–DOE Vera C. Rubin Observatory, first scientific alerts, 25 February 2026 — seven full-stream brokers plus two down-stream services, as its data-products page describes them; alerts world-public [48]

Mathematics — Landmark System, Stated Exactly

Theorem Provers: the Only Field Whose Verifier Cannot Be Fooled


The IMO records the AI companies as a fringe event testing "closed-source AI models… privately" [52]. Benchmarks are not neutral ground either: in June 2026 one research-mathematics benchmark shipped a v2 "addressing errors in 42% of problems". [51]

Section 06

Social Sciences & Humanities

The field where the tool is a general model doing a specialised job — and where accuracy on the label is the easy half of the problem.

Social Sciences — Landmark Result, and the Correction

Annotation at Scale Works. Plugging the Labels Into a Regression Does Not.


"Direct use of surrogate labels in downstream statistical analyses leads to substantial bias and invalid confidence intervals, even with high surrogate accuracy of 80–90%." Egami, Hinck, Stewart & Wei, NeurIPS 2023 — because model errors are non-random and correlate with the covariates you are regressing on [54]

Social Sciences — Where the Story Is Contested

"Silicon Samples": a Debate Fought on 2022–24 Models — and Still Unsettled


The claim

  • GPT-3, 2023. Conditioned on sociodemographic backstories, a model can "accurately emulate response distributions from a wide variety of human subgroups" — the authors' term: algorithmic fidelity [57]
  • Preprint, revised June 2026. Self-report grounding works: across 1,052 Americans, interview-only, survey-only and combined agents reached 83%, 82% and 86% of participants' own two-week test–retest consistency, vs 74% demographics-only [60]

The rebuttals

  • GPT-3.5, 2024. Means match; the rest does not. "Less variation in responses than in the real surveys, and regression coefficients often differ significantly"; "the same prompt yields significantly different results over a 3-month period" [58]
  • Independent replication, 2025. Models "cannot replace research subjects", showing "strong bias and a low variance on each topic" — and "this bias randomly varies from one topic to the next" [59]

Read this pair as a live argument, not a verdict: every number was measured on a 2022–24 model and nobody has re-run either side on the 2026 generation. What survives is structural: matching the mean is not matching the distribution, social science lives in the variance, and a topic-random bias [59] cannot be corrected without the human data you were avoiding. [60] is still an unpublished preprint (v3, checked 2026-09-02) whose revision shows a structured survey (82%) buys what a two-hour interview (83%) does.

Humanities — The Best-Verified Result in This Deck

A Scroll Nobody Could Open, Read End to End


"PHerc. 1667, sealed since the eruption of Vesuvius in 79 AD, has been virtually unwrapped and read from beginning to end… roughly 1.4 metres of papyrus and around twenty-two columns of Greek." Vesuvius Challenge, 25 June 2026 — X-ray tomography, geometric reconstruction, ML ink detection; every reading transcribed and reviewed by papyrologists; data openly licensed, code on GitHub, preprint posted [61]

Series Artefact · Today's Canonical Workflow

The Field-Frontier Tracker


Field Landmark system — and its own stated limit The external verifier Where the frontier is published
Medicineclinical research Imaging classifiers and triage; one autonomous screening exemplar. Limit: the regulator's list is "not comprehensive" [7] [8] Premarket review, then a prospective study in your own setting [10] FDA's device list, updated in place [7] · NEJM AI [65] · your specialty's registry
Biotechlife sciences Joint structure prediction; generative design. Limit: static structures, 4.4% chirality violations, weak ranking [19] [21] CASP, blinded [66] — then 96 wells: binder success 19% [25] or 10–100% [26] CASP assessment papers [66] [21] · AlphaFold DB release notes [22] · bioRxiv, then the journal
Engineering Neural PDE surrogates; an operational ML weather model. Limit: 5–15% in skill; orders of magnitude only in cost [36] [33] A strong classical baseline on an open dataset — 79% of comparisons used a weak one [32] Open benchmark suites such as The Well [38] · ECMWF's AIFS pages, which publish the failures too [36] [37]
Physical sciencesand mathematics Crystal-stability screening; ML alert brokers; RL provers in Lean. Limit: proposed compounds are not materials [44] [49] Synthesis and diffraction re-examined by outsiders — and a journal Author Correction [45] [46]; a live leaderboard [47]; a proof kernel [50] Matbench Discovery [47] · Rubin broker docs [48] · Mathlib and the Lean Zulip [50] · benchmark changelogs [51]
Social sciencesand humanities LLM annotation and coding; HTR on historical hands. Limit: κ collapses on interpretive constructs [56] [58] Human double-coding on a random subset, sampling probability fixed in advance [54] Sociological Methods & Research · Political Analysis [58] [59] · J. Open Humanities Data [63] · SICSS [67]
Columns 3 and 4 are the whole artefact: a field's frontier is not a list of tools — it is a verifier, and a place where results get contested in public. Rebuild it for your own field; it stays true when every product name on it changes.

Series Artefact — Session 5's Canvas, Row 7

The Last Row. The Canvas Is Finished.


Lifecycle stage What you hand to AI Tier · Zone Your non-negotiable check Artefact that governs it
1–5Sessions 1–5 Frame · Discover · Read · Analyse · Write Tier 1–2 · mostly Unreliable As recorded on your canvas CRIT · Matrix · Triage · Guardrails · Policy Table
6 · Cross-fieldSession 6 Field B's vocabulary, landscape and reading list — never its conclusions Tier 1Tier 2 · Unreliable, Dangerous at rung 4 Glossary checked against one field-B review · every paper resolved · 2 of 8 read in full · one field-B reader on the memo The Cross-Field Orientation Ladder
7 · My field's toolsSession 7 — today One bounded, checkable object: a predicted structure, a candidate list, a coded transcript, a formal proof obligation. Never the finding. Outside Tier 1–3: a domain model with a domain verifier · Safe only where the verifier is external and you ran it Name the verifier before you run the tool — experiment · blind benchmark · proof kernel · second human coder — then answer the five questions and record the model version and date The Field-Frontier Tracker
Three rules for row 7. (1) If you cannot name the verifier, the tool is not ready for your project — it is ready for your reading list. (2) A benchmark number is not a deployment number: 87.3% became 0.33 on real patients [16] [17]. (3) Re-date the row every six months — inputs drift, and fine-tuned models lose skill when they do [37].

Demo Preview — You Choose Two, Live

Five Field Demos Are Loaded. We Run the Two You Vote For.


Each runs about five minutes. All five are written out in full in the demo script with prompts, expected outcomes and fallbacks, so the three we do not run are still yours. No real patient, participant or unpublished material appears in any of them — the sample data is synthetic and printed in the script.

Closing the Series

Seven Sessions, Seven Artefacts, One Habit


"The proliferation of AI tools in science risks introducing a phase of scientific enquiry in which we produce more but understand less." Messeri & Crockett, Nature 2024 — on illusions of understanding, and the scientific monocultures they hide [4]

References · Cross-cutting · Medicine · Biotech

References (1–23)


  1. The Royal Swedish Academy of Sciences (2024). The Nobel Prize in Chemistry 2024. nobelprize.org/prizes/chemistry/2024/summary — accessed 2026-07-29
  2. The Royal Swedish Academy of Sciences (2024). The Nobel Prize in Physics 2024. nobelprize.org/prizes/physics/2024/summary — accessed 2026-07-29
  3. Wang, H., et al. (2023). Scientific discovery in the age of artificial intelligence. Nature 620, 47–60. doi.org/10.1038/s41586-023-06221-2 — accessed 2026-07-29
  4. Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature 627, 49–58. doi.org/10.1038/s41586-024-07146-0 — accessed 2026-07-29
  5. Gao, J., & Wang, D. (2024). Quantifying the use and potential benefits of AI in scientific research. Nature Human Behaviour 8, 2281–2292. doi.org/10.1038/s41562-024-02020-5 — accessed 2026-07-29
  6. The Royal Society (2024). Science in the age of AI. royalsociety.org/news-resources/projects/science-in-the-age-of-ai — accessed 2026-07-29
  7. U.S. Food and Drug Administration (2026). Artificial Intelligence-Enabled Medical Devices (list + downloadable file). Content current 2026-06-16. fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices — accessed 2026-09-02
  8. U.S. Food and Drug Administration (2018). De Novo Decision Summary, DEN180001 (IDx-DR). accessdata.fda.gov/cdrh_docs/reviews/DEN180001.pdf — accessed 2026-07-29
  9. Abràmoff, M. D., et al. (2018). Pivotal trial of an autonomous AI-based diagnostic system… npj Digital Medicine 1, 39. doi.org/10.1038/s41746-018-0040-6 — accessed 2026-07-29
  10. Wang, Z., et al. (2025). Systematic review and meta-analysis of regulator-approved deep learning systems for fundus DR detection. npj Digital Medicine. doi.org/10.1038/s41746-025-02223-8 — accessed 2026-07-29
  11. Antonissen, N., et al. (2026). AI in radiology: 173 commercially available products and their scientific evidence. European Radiology. doi.org/10.1007/s00330-025-11830-8 — accessed 2026-07-29
  12. Lawrence, R., et al. (2025). AI for diagnostics in radiology practice: a rapid systematic scoping review. eClinicalMedicine 83, 103228. pubmed.ncbi.nlm.nih.gov/40474995 — accessed 2026-07-29
  13. Goh, E., et al. (2024). Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open 7(10), e2440969. doi.org/10.1001/jamanetworkopen.2024.40969 — accessed 2026-07-29
  14. Goh, E., et al. (2025). GPT-4 assistance for improvement of physician performance on patient care tasks: an RCT. Nature Medicine. doi.org/10.1038/s41591-024-03456-y — accessed 2026-07-29
  15. Qazi, I. A., Ali, A., Khawaja, A. U., et al. (2026). Automation Bias in LLM-Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — An RCT. NEJM AI 3(5). doi.org/10.1056/AIoa2501001 — accessed 2026-07-29
  16. Jin, Q., et al. (2024). Matching patients to clinical trials with large language models (TrialGPT). Nature Communications 15, 9074. doi.org/10.1038/s41467-024-53081-z — accessed 2026-07-29
  17. Kempf, E., et al. (2025). A prospective pragmatic evaluation of automatic trial matching tools in a molecular tumor board. npj Precision Oncology. doi.org/10.1038/s41698-025-00806-y — accessed 2026-07-29
  18. Wei, C.-H., et al. (2024). PubTator 3.0: an AI-powered literature resource… Nucleic Acids Research 52(W1), W540–W546. doi.org/10.1093/nar/gkae235 — accessed 2026-07-29
  19. Abramson, J., et al. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500. doi.org/10.1038/s41586-024-07487-w — accessed 2026-07-29
  20. Google DeepMind (2026). AlphaFold Server FAQ; alphafold3 repository; Model Parameters Terms of Use. alphafoldserver.com/faq (2026-07-29) · github.com/google-deepmind/alphafold3 — accessed 2026-09-02
  21. Yuan, R., et al. (2025). CASP16 Protein Monomer Structure Prediction Assessment. Proteins. doi.org/10.1002/prot.70031 — accessed 2026-07-29
  22. EMBL-EBI & Google DeepMind (2026). AlphaFold Protein Structure Database; with Varadi, M., et al. (2024). NAR 52, D368–D375. alphafold.ebi.ac.uk · doi.org/10.1093/nar/gkad1011 — accessed 2026-07-29
  23. Peng, C., et al. (2025). A comprehensive benchmarking of AlphaFold3… Briefings in Bioinformatics 26(6), bbaf616. doi.org/10.1093/bib/bbaf616 — accessed 2026-07-29

References · Biotech · Engineering · Physical Sciences

References (24–45)


  1. Dauparas, J., et al. (2022). Robust deep learning–based protein sequence design using ProteinMPNN. Science 378, 49–56. doi.org/10.1126/science.add2187 — accessed 2026-07-29
  2. Watson, J. L., et al. (2023). De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100. doi.org/10.1038/s41586-023-06415-8 — accessed 2026-07-29
  3. Pacesa, M., et al. (2025). One-shot design of functional protein binders with BindCraft. Nature 646(8084), 483–492. doi.org/10.1038/s41586-025-09429-6 — accessed 2026-07-29
  4. Lauko, A., Ahern, W., et al. (2026). Atom-level enzyme active site scaffolding using RFdiffusion2. Nature Methods 23, 96–105. doi.org/10.1038/s41592-025-02975-x — accessed 2026-07-29
  5. Hayes, T., et al. (2025). Simulating 500 million years of evolution with a language model. Science 387, eads0018. doi.org/10.1126/science.ads0018 — accessed 2026-07-29
  6. Jayatunga, M. K. P., et al. (2024). How successful are AI-discovered drugs in clinical trials? Drug Discovery Today 29(6), 104009. doi.org/10.1016/j.drudis.2024.104009 — accessed 2026-07-29
  7. Insilico Medicine, et al. (2025). A generative AI-discovered TNIK inhibitor for IPF: a randomized phase 2a trial. Nature Medicine. doi.org/10.1038/s41591-025-03743-2 — accessed 2026-07-29
  8. Kedzierska, K. Z., et al. (2025). Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biology 26, 101. doi.org/10.1186/s13059-025-03574-x — accessed 2026-07-29
  9. McGreivy, N., & Hakim, A. (2024). Weak baselines and reporting biases lead to overoptimism in ML for fluid-related PDEs. Nature Machine Intelligence 6. doi.org/10.1038/s42256-024-00897-5 · arxiv.org/abs/2407.07218 — accessed 2026-07-29
  10. Krishnapriyan, A. S., et al. (2021). Characterizing possible failure modes in physics-informed neural networks. NeurIPS 2021. arxiv.org/abs/2109.01050 — accessed 2026-07-29
  11. Azizzadenesheli, K., et al. (2024). Neural operators for accelerating scientific simulations and design. Nature Reviews Physics 6, 320–328. doi.org/10.1038/s42254-024-00712-5 — accessed 2026-07-29
  12. Lam, R., et al. (2023). Learning skillful medium-range global weather forecasting (GraphCast). Science 382, eadi2336. doi.org/10.1126/science.adi2336 — accessed 2026-07-29
  13. ECMWF (2026). AIFS Machine Learning data; and Haiden & Chevallier (2026), Forecast performance 2025, Newsletter 187. ecmwf.int/en/forecasts/dataset/aifs-machine-learning-data — accessed 2026-07-29
  14. Ben Bouallègue, Z., Raoult, B., & Chantry, M. (2026). Farewell to the external AI models. ECMWF AIFS Blog, 11 May 2026. doi.org/10.21957/fa8ad01483 — accessed 2026-07-29
  15. Ohana, R., et al. (2024). The Well: a large-scale collection of diverse physics simulations for ML. NeurIPS 2024 D&B. arxiv.org/abs/2412.00568 · polymathic-ai.org/the_well — accessed 2026-07-29
  16. Autodesk (2020; product page 2026). Topology Optimization is not Generative Design; Generative design for manufacturing. autodesk.com/products/fusion-360/blog · autodesk.com/solutions/generative-design/manufacturing — accessed 2026-07-29
  17. Zoo (2026). Frequently Asked Questions; Text-to-CAD. zoo.dev/docs/faq · zoo.dev/design-studio (v1.4.4) — accessed 2026-09-02
  18. Pearce, H., et al. (2022). Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions. IEEE S&P 2022, 754–768. doi.org/10.1109/SP46214.2022.9833571 · arxiv.org/abs/2108.09293 — accessed 2026-07-29
  19. Becker, J., et al. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. Preprint. arxiv.org/abs/2507.09089 — accessed 2026-07-29
  20. Merchant, A., et al. (2023). Scaling deep learning for materials discovery (GNoME). Nature 624, 80–90. doi.org/10.1038/s41586-023-06735-9 — accessed 2026-07-29
  21. Cheetham, A. K., & Seshadri, R. (2024). Artificial Intelligence Driving Materials Discovery? Chemistry of Materials 36(8), 3490–3495. doi.org/10.1021/acs.chemmater.4c00643 — accessed 2026-07-29
  22. Szymanski, N. J., et al. (2023). An autonomous laboratory for the accelerated synthesis of inorganic materials (A-Lab). Nature 624, 86–91. Read with [69]. doi.org/10.1038/s41586-023-06734-w — accessed 2026-07-29

References · Physical Sciences & Mathematics · Social Sciences

References (46–58)


  1. Leeman, J., et al. (2024). Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. PRX Energy 3, 011002. doi.org/10.1103/PRXEnergy.3.011002 — accessed 2026-07-29
  2. Riebesell, J., et al. (2025). A framework to evaluate machine learning crystal stability predictions (Matbench Discovery). Nature Machine Intelligence 7, 836–847. doi.org/10.1038/s42256-025-01055-1 · matbench-discovery.materialsproject.org — accessed 2026-09-02
  3. NSF–DOE Vera C. Rubin Observatory (2026). Rubin Observatory Launches Real-Time Alerts (25 Feb 2026); Alerts and brokers. rubinobservatory.org/news/first-alerts — accessed 2026-07-29
  4. Hubert, T., Masoom, R., Barekatain, M., et al. / Google DeepMind (2025). Olympiad-level formal mathematical reasoning with reinforcement learning (AlphaProof). Nature. doi.org/10.1038/s41586-025-09833-y — accessed 2026-07-29
  5. Lean community (2026). Lean 4 Web; LeanSearch; Mathlib statistics. live.lean-lang.org · leansearch.net · leanprover-community.github.io/mathlib_stats.html — accessed 2026-09-02
  6. Epoch AI (2026). FrontierMath; FrontierMath Tiers 1–4 (v2, 12 June 2026). epoch.ai/frontiermath — accessed 2026-07-29
  7. IMO 2025 Organisers / Dolinar, G. (2025). The 66th International Mathematical Olympiad draws to a close today, 19 July 2025. imo2025.au/news/the-66th-international-mathematical-olympiad-draws-to-a-close-today — accessed 2026-07-29
  8. Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. PNAS 120(30), e2305016120. doi.org/10.1073/pnas.2305016120 — accessed 2026-07-29
  9. Egami, N., Hinck, M., Stewart, B. M., & Wei, H. (2023). Using Imperfect Surrogates for Downstream Inference (DSL). NeurIPS 2023. arxiv.org/abs/2306.04746 · naokiegami.com/dsl — accessed 2026-07-29
  10. Ollion, É., Shen, R., Macanovic, A., & Chatelain, A. (2024). The dangers of using proprietary LLMs for research. Nature Machine Intelligence 6, 4–5. doi.org/10.1038/s42256-023-00783-6 — accessed 2026-07-29
  11. Liu, X., et al. (2025). Qualitative Coding with GPT-4: Where it Works Better. J. Learning Analytics 12(1), 169–185. doi.org/10.18608/jla.2025.8575 — accessed 2026-07-29
  12. Argyle, L. P., et al. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31(3), 337–351. doi.org/10.1017/pan.2023.2 — accessed 2026-07-29
  13. Bisbee, J., et al. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis 32(4), 401–416. doi.org/10.1017/pan.2024.5 — accessed 2026-07-29

References · Social Sciences & Humanities · Frontier Venues

References (59–70)


  1. Boelaert, J., et al. (2025). Machine Bias: How Do Generative Language Models Answer Opinion Polls? Sociological Methods & Research. doi.org/10.1177/00491241251330582 — accessed 2026-07-29
  2. Park, J. S., et al. (2026). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals. Preprint, v3, never journal-published. arxiv.org/abs/2411.10109 — accessed 2026-09-02
  3. Vesuvius Challenge (2026). An entire Herculaneum scroll has been read for the first time, 25 June 2026. scrollprize.org/firstscroll · scrollprize.org/data — accessed 2026-07-29
  4. READ-COOP SCE / Transkribus (2026). Pricing; Character Error Rate (CER) explained. transkribus.org/pricing · transkribus.org/character-error-rate-cer-explained — accessed 2026-07-29
  5. Hutchinson, D. (2026). Benchmarking as Source Criticism: From Recognition to Reasoning in LLM Assessment. J. Open Humanities Data 12, art. 43. doi.org/10.5334/johd.489 — accessed 2026-07-29
  6. MAXQDA/VERBI, Lumivero (NVivo), ATLAS.ti, Taguette (2026). AI feature documentation. maxqda.com/products/ai-assist · lumivero.com/products/nvivo · atlasti.com · taguette.org/about.html — accessed 2026-07-29
  7. NEJM Group (2026). About NEJM AI. ISSN 2836-9386. ai.nejm.org/about — accessed 2026-07-29
  8. Protein Structure Prediction Center (2026). CASP16. predictioncenter.org/casp16 — accessed 2026-07-29
  9. Summer Institutes in Computational Social Science (2026). About SICSS. sicss.io/about — accessed 2026-07-29
  10. Jacobson, R. D. (2025). The AI drug revolution needs a revolution. npj Drug Discovery. doi.org/10.1038/s44386-025-00013-6 — accessed 2026-07-29
  11. Szymanski, N. J., et al. (2026). Author Correction: An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature 650, E1. doi.org/10.1038/s41586-025-09992-y — accessed 2026-07-29
  12. Kim, D., Woodbury, S. M., Ahern, W., et al. (2025). Computational design of metallohydrolases. Nature 649, 246–253. doi.org/10.1038/s41586-025-09746-w — accessed 2026-08-08