← series index
# Sources — Session 6: AI for Multidisciplinary Research
*AI for Researchers · dated, annotated reference list · compiled 2026-07-29 · revised 2026-08-21 · Landscape as of August 2026*
Entry format (one line per source, numbering matches the `[n]` footnote markers used in `slides.html` and `handout.md`):
`- [n] Author/Org (Year). *Title*. URL — accessed YYYY-MM-DD. [type: peer-reviewed | policy | primary-doc | news] — one-line note on what claim(s) it grounds.`
Allowed `type` values:
- `peer-reviewed` — peer-reviewed paper or arXiv/preprint literature
- `policy` — official publisher/funder/institutional policy page
- `primary-doc` — primary tool/vendor documentation (not a blog or listicle)
- `news` — recent news/benchmark article; informs tool discovery only, never the sole ground for a factual claim
**Evidence rule for this session.** The claim "crossing fields is hard" is the load-bearing claim of the whole deck, so it is grounded exclusively in the bibliometrics, science-of-science and team-science literature — never in assertion or anecdote. Every number about *how well AI performs* comes from a published evaluation with a stated task set. Every fact about *what a tool does* comes from that tool's own current documentation and nowhere else; where a vendor does not document something, this deck says "not documented" rather than guessing. Preprints are labelled as preprints on the slide that uses them.
---
## References
### A. Why crossing fields is hard — the interdisciplinarity, bibliometrics and team-science literature
- [1] Uzzi, B., Mukherjee, S., Stringer, M., & Jones, B. (2013). *Atypical Combinations and Scientific Impact*. Science 342(6157), 468–472. https://doi.org/10.1126/science.1240474 — accessed 2026-07-29. [type: peer-reviewed] — Grounds the opening "prize" slide. Across **17.9 million papers** spanning all scientific fields, the highest-impact work combines exceptionally conventional prior work with an intrusion of unusual combinations; quoted verbatim on the slide: "Papers of this type were **twice as likely to be highly cited** works. Novel combinations of prior work are rare, yet teams are **37.7% more likely** than solo authors to insert novel combinations into familiar knowledge domains." (science.org returns 403 to automated fetch; abstract verified via PubMed 24159044 and the DOI verified live in Crossref on 2026-07-29.)
- [2] Larivière, V., & Gingras, Y. (2010). *On the relationship between interdisciplinarity and scientific impact*. Journal of the American Society for Information Science and Technology 61(1), 126–131. https://doi.org/10.1002/asi.21226 (open preprint: https://arxiv.org/abs/0908.1776) — accessed 2026-07-29. [type: peer-reviewed] — The inverted-U shape of interdisciplinarity and impact, measured on all Web of Science papers published in 2000. Quoted: "highly disciplinary and highly interdisciplinary papers have a low scientific impact… there might be an **optimum of interdisciplinarity** beyond which the research is too dispersed to find its niche and under which it is too mainstream to have high impact." Verified against the open arXiv preprint text; the Wiley version is paywalled.
- [3] Yegros-Yegros, A., Rafols, I., & D'Este, P. (2015). *Does Interdisciplinary Research Lead to Higher Citation Impact? The Different Effect of Proximal and Distal Interdisciplinarity*. PLOS ONE 10(8), e0135095. https://doi.org/10.1371/journal.pone.0135095 — accessed 2026-07-29. [type: peer-reviewed] — Decomposes interdisciplinarity into Stirling's variety, balance and disparity [12] and shows they pull in *opposite* directions: "variety has a positive effect on impact, whereas balance and disparity have a negative effect… all three dimensions of interdisciplinarity display a **curvilinear (inverted U-shape)** relationship with citation impact." Grounds the slide's claim that "interdisciplinary" is not one thing.
- [4] Wang, J., Thijs, B., & Glänzel, W. (2015). *Interdisciplinarity and Impact: Distinct Effects of Variety, Balance, and Disparity*. PLOS ONE 10(5), e0127298. https://doi.org/10.1371/journal.pone.0127298 — accessed 2026-07-29. [type: peer-reviewed] — The delayed-recognition finding, quoted on the slide: "although variety and disparity have positive effects on long-term citations, they have **negative effects on short-term (3-year) citations**." Long-term window is 13 years. **Not used:** per-standard-deviation percentages surfaced during research could not be traced to the results tables, so no such figure appears anywhere in this deck.
- [5] Leahey, E., Beckman, C. M., & Stanko, T. L. (2017). *Prominent but Less Productive: The Impact of Interdisciplinarity on Scientists' Research*. Administrative Science Quarterly 62(1), 105–139. https://doi.org/10.1177/0001839216665364 — accessed 2026-07-29. [type: peer-reviewed] — Almost **900 research-centre scientists** and their **32,000 articles**, including unpublished work. Quoted: interdisciplinarity carries "both penalties (fewer papers published) and benefits (increased citations)… it is a **high-risk, high-reward endeavor**." The abstract states no numeric effect sizes, so the deck states none.
- [6] Bromham, L., Dinnage, R., & Hua, X. (2016). *Interdisciplinary research has consistently lower funding success*. Nature 534(7609), 684–687. https://doi.org/10.1038/nature18315 — accessed 2026-07-29. [type: peer-reviewed] — Quoted: "Using data on **all 18,476 proposals** submitted to the scheme over 5 consecutive years, including successful and unsuccessful applications, we show that the **greater the degree of interdisciplinarity, the lower the probability of being funded**." Significant after controlling for number of collaborators, primary field and institution type. **Explicit non-claim:** the abstract reports a direction, not a percentage-point gap, so the deck quotes no "X% lower" figure. (nature.com 403s to automated fetch; abstract verified via Europe PMC, DOI verified in Crossref 2026-07-29.)
- [7] Berkes, E., Marion, M., Milojević, S., & Weinberg, B. A. (2024). *Slow convergence: Career impediments to interdisciplinary biomedical research*. PNAS 121(32), e2402646121. https://doi.org/10.1073/pnas.2402646121 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC11317606/) — accessed 2026-07-29. [type: peer-reviewed] — The deck's most arresting single statistic, quoted verbatim: across **154,021 biomedical PhDs, 1970–2013**, "it takes about **8 y for half of the researchers in the top percentile in terms of initial interdisciplinarity to stop publishing**, compared to **more than 20 y for moderately interdisciplinary researchers** (10th to 75th percentiles)." Also grounds the slide's line that initially interdisciplinary researchers *reduce* their interdisciplinarity over time.
- [8] Xiang, S., Romero, D. M., & Teplitskiy, M. (2025). *Evaluating interdisciplinary research: Disparate outcomes for topic and knowledge base*. PNAS 122(17), e2409752122. https://doi.org/10.1073/pnas.2409752122 — accessed 2026-07-29. [type: peer-reviewed] — Peer-review outcomes for **128,950 manuscripts including rejections**. Quoted: "higher **knowledge-base** interdisciplinarity (measured through references) was associated with **higher** acceptance rates, and higher **topic** interdisciplinarity (measured through title and abstract text) was associated with **lower** ones." This is the deck's practical advice on how to present cross-field work — verified from the Significance statement via the Semantic Scholar record; pnas.org 403s to automated fetch.
- [9] Martínez, A., & Mammola, S. (2021). *Specialized terminology reduces the number of citations of scientific papers*. Proceedings of the Royal Society B 288(1948), 20202581. https://doi.org/10.1098/rspb.2020.2581 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC8059506/) — accessed 2026-07-29. [type: peer-reviewed] — **21,486 articles** in cave research, chosen as "a multidisciplinary field particularly prone to terminological specialization". Quoted: "in the era of interdisciplinarity, the use of jargon may **hinder effective communication among scientists that do not share a common scientific background**… We demonstrate a significant negative relationship between the proportion of jargon words in the title and abstract and the number of citations a paper receives." Grounds the slide's claim that vocabulary is a measurable, not merely felt, barrier.
- [10] Vilhena, D. A., Foster, J. G., Rosvall, M., West, J. D., Evans, J., & Bergstrom, C. T. (2014). *Finding Cultural Holes: How Structure and Culture Diverge in Networks of Scholarly Communication*. Sociological Science 1, 221–238. https://doi.org/10.15195/v1.a15 — accessed 2026-07-29. [type: peer-reviewed] — Measures the *communicative* distance between fields from phrase frequencies in JSTOR full text, independently of the citation graph. Quoted: "the ecological sciences are **balkanized by jargon**, whereas the social sciences are relatively integrated." Grounds the deck's claim that the distance between two fields is field-pair-specific and computable.
- [11] National Research Council (2015). *Enhancing the Effectiveness of Team Science*. Committee on the Science of Team Science; N. J. Cooke & M. L. Hilton (Eds.). National Academies Press. https://doi.org/10.17226/19007 (open full text: https://www.ncbi.nlm.nih.gov/books/NBK310391/) — accessed 2026-07-29. [type: peer-reviewed] — Consensus study report grounding the team-science section. The **seven features that create challenges**, quoted verbatim: "(1) high diversity of membership; (2) deep knowledge integration; (3) large size; (4) goal misalignment with other teams; (5) permeable team and group boundaries; (6) geographic dispersion; and (7) high task interdependence." Mechanism quoted on the slide: "In highly diverse team science projects, **communication problems can occur because of members' use of technical or scientific language that is unique to their area of expertise** and therefore unfamiliar to other members"; and "The unique languages of the disciplines reflect **deeper differences in underlying assumptions, epistemologies (ways of knowing), philosophies**, and approaches to science and societal problems."
- [12] Stirling, A. (2007). *A general framework for analysing diversity in science, technology and society*. Journal of the Royal Society Interface 4(15), 707–719. https://doi.org/10.1098/rsif.2007.0213 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC2373389/) — accessed 2026-07-29. [type: peer-reviewed] — The origin of the variety / balance / disparity framework that [3] and [4] measure, and the reason the deck refuses to treat "interdisciplinary" as a single dial. Definitions quoted: variety = "the number of categories into which system elements are apportioned"; balance = "a function of the pattern of apportionment of elements across categories"; disparity = "the manner and degree in which the elements may be distinguished".
- [13] Wang, Q., & Schneider, J. W. (2020). *Consistency and validity of interdisciplinarity measures*. Quantitative Science Studies 1(1), 239–263. https://doi.org/10.1162/qss_a_00011 — accessed 2026-07-29. [type: peer-reviewed] — The honest caveat placed under the bibliometric evidence: measures that should capture the same dimension give "surprisingly deviant results", and "the current measurements of interdisciplinarity should be **interpreted with much caution**". Cited on the slide so the audience does not over-read [3] and [4].
- [14] McLeish, T., & Strang, V. (2016). *Evaluating interdisciplinary research: the elephant in the peer-reviewers' room*. Palgrave Communications 2, 16055. https://doi.org/10.1057/palcomms.2016.55 — accessed 2026-07-29. [type: peer-reviewed] — Grounds the claim that single-discipline peer review is poorly suited to interdisciplinary work. Quoted: "The difference between high-quality and poor IDR is most often not to be found in the quality of its disciplinary ingredients… but rather in **how these are combined** to generate the whole research project." Conceptual, not empirical — labelled as such on the slide.
- [15] U.S. National Science Foundation (2026). *Learn About Convergence Research* (Research Approaches). https://www.nsf.gov/funding/learn/research-types/learn-about-convergence-research — accessed 2026-07-29. [type: policy] — A funder's own definition, used to show that the thing this session teaches is explicitly what funders now ask for. NSF names two primary characteristics: research "driven by a specific and compelling problem" and "deep integration across disciplines", in which researchers' "knowledge, theories, methods, data and research communities increasingly intermingle" and they "**develop a shared scientific language**".
### B. What AI actually does — and fails to do — across field boundaries
- [16] Hao, Q., Xu, F., Li, Y., & Evans, J. (2026). *Artificial intelligence tools expand scientists' impact but contract science's focus*. Nature 649(8099), 1237–1243. https://doi.org/10.1038/s41586-025-09922-y — accessed 2026-07-29. [type: peer-reviewed] — The deck's central tension, and the reason this session is about *orientation* rather than *delegation*. Analysing **41.3 million papers**: AI-augmented scientists publish **3.02× more papers**, receive **4.84× more citations** and become project leaders **1.37 years earlier** — while AI adoption "shrinks the collective volume of scientific topics studied by **4.63%**" and decreases scientists' engagement with one another by **22%**. Verified via the PubMed record (PMID 41535462) and the Crossref record; nature.com 403s to automated fetch.
- [17] Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., & Bowman, S. R. (2024). *GPQA: A Graduate-Level Google-Proof Q&A Benchmark*. COLM 2024. https://arxiv.org/abs/2311.12022 — accessed 2026-07-29. [type: peer-reviewed] — The best published quantification of how far you fall off a cliff outside your own field, quoted on the slide: on **448** biology/physics/chemistry questions, "experts who have or are pursuing PhDs in the corresponding domains reach **65%** accuracy (74% when discounting clear mistakes…), while **highly skilled non-expert validators only reach 34%** accuracy, despite spending on average **over 30 minutes with unrestricted access to the web**." Used to make a claim about *humans*, not about models.
- [18] Wang, Y., Ma, X., Zhang, G., et al. (2024). *MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark*. NeurIPS 2024 Datasets & Benchmarks Track. https://arxiv.org/abs/2406.01574 — accessed 2026-07-29. [type: peer-reviewed] — **12,032 questions across 14 disciplines**; overall accuracy falls "by 16% to 33% compared to MMLU". The deck's by-discipline spread is read directly from the paper's Table 2 (*Models Performance on MMLU-Pro, CoT*), verified independently on the arXiv HTML: for a single model (GPT-4o) accuracy runs from **Psychology 79.2%** down to **Law 51.0%** — a **28-point gap between disciplines for the same model**. Stated on the slide as a per-model, per-discipline spread, not as a ranking of fields. **Vintage note (2026-08-21 revision):** this spread is a 2024 measurement of GPT-4o, and the slide now labels it "GPT-4o, 2024". By August 2026 frontier models cluster at **83–90%** on MMLU-Pro (Artificial Analysis: Gemini 3 Pro Preview 89.8%, Claude Opus 4.5 (Reasoning) 89.5% [65]) and the benchmark is regarded as near-saturated at the frontier; **no verified 2026 per-discipline spread was found**, so the deck carries the live expert-frontier evidence on HLE ([19], [65]) instead of silently updating this one.
- [19] Phan, L., Gatti, A., Han, Z., Li, N., et al. (2026). *A benchmark of expert-level academic questions to assess AI capabilities* (Humanity's Last Exam). Nature 649(8099), 1139–1146. https://doi.org/10.1038/s41586-025-09962-4 (preprint: https://arxiv.org/abs/2501.14249) — accessed 2026-07-29. [type: peer-reviewed] — **2,500 questions** from ~1,000 expert contributors across 50 countries and **over 100 subjects**. Quoted: "State-of-the-art LLMs demonstrate **low accuracy and calibration** on HLE, highlighting a marked gap between current LLM capabilities and the expert human frontier on closed-ended academic questions." The deck uses the *calibration* point as the durable claim; as of the 2026-08-21 revision it also states one dated leaderboard score — **55.5% (Claude Fable 5)** per Artificial Analysis's independent HLE run [65] — labelled with its source and date precisely *because* leaderboard scores move monthly. Better than the 2025-era single digits, still far below the expert human frontier.
- [20] Kang, Z., Gong, J., Yan, J., Xia, W., Wang, Y., et al. (2026). *HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models*. **Preprint** (v3, March 2026). https://arxiv.org/abs/2506.03922 — accessed 2026-07-29. [type: peer-reviewed] — Labelled as a preprint on the slide. Provides the deck's framing quote, written by the benchmark's own authors: "current benchmarks… primarily emphasize general knowledge and **vertical step-by-step reasoning typical of STEM disciplines**, while overlooking the distinct needs and potential of the Humanities and Social Sciences (HSS). Tasks in the HSS domain require more **horizontal, interdisciplinary thinking** and a deep integration of knowledge across related fields, which presents unique challenges for MLLMs." **13,152 samples**, six categories; best open-ended score 39.97% against a human expert baseline of 93.83%.
- [21] Linardon, J., Jarman, H., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). *Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study*. JMIR Mental Health 12, e80371. https://doi.org/10.2196/80371 — accessed 2026-07-29. [type: peer-reviewed] — The single most important source for this session: fabrication is **not uniform across topics**. Across six literature reviews, 176 citations were produced; **35 (19.9%) were fabricated**, and of the 141 genuine ones **64 (45.4%) contained errors**. Fabrication varied significantly with topic familiarity (χ²₂ = 13.7; *P* = .001): **6% (4/68)** for major depressive disorder against **28% (17/60)** for binge eating disorder and **29% (14/48)** for body dysmorphic disorder. **Stated precisely, because this is the entry's one easy-to-misread number:** the *overall* comparison of specialised against general prompts was **not significant** (*P* = .21); the striking 46% vs 17% contrast (*P* = .01) is the **binge-eating-disorder subgroup** only. No slide, handout or demo line in this session uses either figure — the deck rests on the topic-familiarity gradient, which is the paper's significant finding. This is exactly the risk profile of a researcher working outside their home field. Verified via the PubMed record (PMID 41223407) and Crossref; the journal page rendered blank to automated fetch. **Currency note (2026-08-21 revision):** a 2025 measurement of an ungrounded chatbot workflow. Grounded current-generation assistants fabricate fewer non-existent references, so the deck keeps the 6%→28–29% figures as a **labelled baseline for how rarity raises risk**, not as a current absolute rate; the current-generation companion evidence is [66] (links valid, factual support 39–77%). No re-run of the topic-familiarity design on the August-2026 cohort was found.
- [22] Peters, U., & Chin-Yee, B. (2025). *Generalization bias in large language model summarization of scientific research*. Royal Society Open Science 12(4), 241776. https://doi.org/10.1098/rsos.241776 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC12042776/) — accessed 2026-07-29. [type: peer-reviewed] — Grounds the warning attached to rung 1 of the Ladder. **10 LLMs, 4,900 generated summaries.** Quoted: "Even when explicitly prompted for accuracy, most LLMs produced **broader generalizations** of scientific results than those in the original texts, with DeepSeek, ChatGPT-4o and LLaMA 3.3 70B **overgeneralizing in 26–73% of cases**." Against human-authored summaries the odds ratio was **4.85 (95% CI 3.06–7.70, *P* < 0.001)**. Two counterintuitive findings the deck states out loud: prompting for accuracy roughly **doubled** overgeneralizations, and "newer models tended to perform worse in generalization accuracy than earlier ones." **Cohort note (2026-08-21 revision):** the ten models are the 2024–25 cohort, so "newer models performed worse" is a **within-cohort** observation, and the slide now says so. The same team's 2026 follow-up [64] extends the finding to GPT-5-class models (ChatGPT-5, DeepSeek-V3.1): both rated generic claims as *more* generalizable than laypeople while domain experts rated them *less*. No overgeneralization study of the August-2026 cohort (GPT-5.6, Claude Fable 5, Kimi K3, Gemini 3.x) exists yet — that gap is stated rather than papered over.
- [23] Guo, Y., Sohn, J. H., Leroy, G., & Cohen, T. (2026). *Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation*. Journal of Biomedical Informatics 179, 105038. https://doi.org/10.1016/j.jbi.2026.105038 (preprint: https://arxiv.org/abs/2505.10409) — accessed 2026-07-29. [type: peer-reviewed] — **150 participants**, subjective ratings plus objective comprehension and recall tests. Quoted: "while LLMs can generate PLSs that appear **indistinguishable from human-written ones in subjective evaluations**, **human-written PLSs lead to significantly better comprehension**"; and "automated evaluation metrics fail to reflect human judgment". The deck uses this to separate *feeling oriented* from *being oriented* — the central risk of rung 1. **Cohort note (2026-08-21 revision):** the preprint dates to May 2025, so the summaries come from pre-2026 models; the slide attribution now says "pre-2026 models".
- [24] Goldsack, T., Scarton, C., Shardlow, M., & Lin, C. (2024). *Overview of the BioLaySumm 2024 Shared Task on the Lay Summarization of Biomedical Research Articles*. Proceedings of the 23rd Workshop on Biomedical Natural Language Processing (BioNLP @ ACL 2024). https://aclanthology.org/2024.bionlp-1.10/ — accessed 2026-07-29. [type: peer-reviewed] — **53 participating teams**, scored on Relevance, Readability and Factuality. Grounds the deck's claim that readability and faithfulness trade off against each other rather than improving together: the organisers report "a **trade-off between scoring highly for Factuality and the metrics of Relevance or Readability**", and that "systems that ranked the highest were those that most successfully balanced this trade-off". Competitive shared task rather than a single lab, which is why it is used; the slide labels it a 2024 shared task, and the trade-off is presented as a task-level finding rather than a claim about current systems.
- [25] Lewis, M., & Mitchell, M. (2024). *Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models*. **Preprint**. https://arxiv.org/abs/2402.08955 (companion: *Evaluating the Robustness of Analogical Reasoning in Large Language Models*, https://arxiv.org/abs/2411.14215) — accessed 2026-07-29. [type: peer-reviewed] — Labelled as a preprint on the slide. Analogical transfer is the core cognitive move of cross-field work, and this is the best-controlled evidence that LLM analogy-making tracks surface similarity: on counterfactual variants that "test the same abstract reasoning abilities but are likely dissimilar from tasks in the pre-training data", **humans scored 75.3%** against **GPT-4 at 45.2%** and GPT-3.5 at 35.0%; human performance "stays relatively constant" while "the GPT models' performance declines sharply". Model vintage (GPT-3.5/GPT-4, early 2024) is stated on the slide.
- [26] Shen, Y., de Sousa, D. X., Marçal, R., Guo, H., & Zhu, X. (2026). *IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research*. **Preprint** (v3, June 2026). https://arxiv.org/abs/2507.15736 — accessed 2026-07-29. [type: peer-reviewed] — The only benchmark found that targets interdisciplinary research directly; labelled as a preprint. Three tasks (IDR Paper Identification, Idea Integration, Idea Recommendation) over ten mainstream LLMs. On identification, best Macro-F1 was **0.640** against 0.176 for random guessing on a **3,675-instance** set built at the natural 1:10 sparsity of interdisciplinary papers. Authors' conclusion, quoted: "LLMs still struggle to reliably distinguish true interdisciplinary integration, and the reasoning-oriented models could degrade IDR performance." **Cohort note (2026-08-21 revision):** the ten models are a pre-current-generation (2025) cohort, evaluated before adaptive/extended thinking became the default — the "reasoning-oriented models could degrade IDR" clause is exactly the claim the August-2026 generation is built to contest, and no re-run on that generation exists. The slide labels the 0.640 figure "2025 cohort".
- [27] Wright, D., Masud, S., Moore, J., Yadav, S., Antoniak, M., Christensen, P. E., Park, C. Y., & Augenstein, I. (2026). *Epistemic Diversity and Knowledge Collapse in Large Language Models*. **Preprint** (v6, January 2026). https://arxiv.org/abs/2510.04226 — accessed 2026-07-29. [type: peer-reviewed] — Labelled as a preprint. **27 LLMs, 155 topics, 200 prompt templates** sourced from real user chats. The finding used on the slide: "while newer models tend to generate more diverse claims, **all models are less epistemically diverse than a basic web search**"; model size has a *negative* effect on diversity, and retrieval-augmented generation has a positive one. The deck's justification for not letting a chatbot alone define an unfamiliar field's landscape. **Cohort note (2026-08-21 revision):** the 27 models run through the 2025 generation; no epistemic-diversity measurement of the August-2026 cohort was found, and the slide and handout now say so. The size-diversity and retrieval findings are stated as that cohort's results, not as laws.
- [28] Dathe, A., Hoffmann, K., & Mangold, A. (2026). *Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research*. **Preprint** (v2, May 2026). https://arxiv.org/abs/2605.10125 — accessed 2026-07-29. [type: peer-reviewed] — Labelled as a preprint. Benchmarks five literature-search tools. Precision (usable ÷ retrieved sources) ranged **21.1%–41.2%**; run-to-run reproducibility, measured as the Jaccard overlap of results for repeated identical queries, ranged **11.8%–28%**. The deck uses the reproducibility number to justify rung 2's obligation to rebuild the landscape in a second, structurally different tool.
### C. Different fields, different evidence standards — the sources behind the Compatibility Check
- [29] Benjamin, D. J., Berger, J. O., Johannesson, M., et al. (2018). *Redefine statistical significance*. Nature Human Behaviour 2(1), 6–10. https://doi.org/10.1038/s41562-017-0189-z — accessed 2026-07-29. [type: peer-reviewed] — One source carries the whole "different thresholds" slide, because the authors make the cross-field comparison themselves. Quoted verbatim: "Recognition of this issue led the genetics research community to move to a '**genome-wide significance threshold' of 5 × 10⁻⁸** over a decade ago. And in high-energy physics, the tradition has long been to define significance by a '**5-sigma' rule (roughly a P value threshold of 3 × 10⁻⁷**). We are essentially suggesting a move from a 2-sigma rule to a 3-sigma rule." The proposal itself: **P < 0.005** for claims of new discoveries, with 0.005 < P < 0.05 relabelled "suggestive". **72 authors** — counted independently on the publisher PDF byline and the PubMed author list, which agree. (nature.com is behind an auth redirect; verified on the publisher PDF at imai.fas.harvard.edu and PubMed 30980045.)
- [30] Wasserstein, R. L., & Lazar, N. A. (2016). *The ASA Statement on p-Values: Context, Process, and Purpose*. The American Statistician 70(2), 129–133. https://doi.org/10.1080/00031305.2016.1154108 (official PDF: https://www.amstat.org/asa/files/pdfs/P-ValueStatement.pdf) — accessed 2026-07-29. [type: policy] — The professional body's own six principles, read verbatim off the ASA's official PDF (the journal page 403s to automated fetch; the DOI itself was verified live in Crossref). Principle 3 is the one on the slide: "Scientific conclusions and business or policy decisions **should not be based only on whether a p-value passes a specific threshold**." Also quoted, from ASA president Jessica Utts on the same page, because it is the cross-field point in one line: "the p-value has become a gatekeeper for whether work is publishable, **at least in some fields**."
- [31] Open Science Collaboration (2015). *Estimating the reproducibility of psychological science*. Science 349(6251), aac4716. https://doi.org/10.1126/science.aac4716 — accessed 2026-07-29. [type: peer-reviewed] — **100 studies.** Quoted: "**Ninety-seven percent** of original studies had statistically significant results. **Thirty-six percent** of replications had statistically significant results"; "Replication effects were **half the magnitude** of original effects." Verified via PubMed 26315443 and Crossref.
- [32] Camerer, C. F., Dreber, A., Forsell, E., et al. (2016). *Evaluating replicability of laboratory experiments in economics*. Science 351(6280), 1433–1436. https://doi.org/10.1126/science.aaf0918 — accessed 2026-07-29. [type: peer-reviewed] — **18 studies** from AER and QJE, all replications ≥90% power. Quoted: "We found a significant effect in the same direction as in the original study for **11 replications (61%)**; on average, the replicated effect size is **66%** of the original."
- [33] Camerer, C. F., Dreber, A., Holzmeister, F., et al. (2018). *Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015*. Nature Human Behaviour 2, 637–644. https://doi.org/10.1038/s41562-018-0399-z — accessed 2026-07-29. [type: peer-reviewed] — **21 studies.** Quoted: "We find a significant effect in the same direction as the original study for **13 (62%) studies**, and the effect size of the replications is on average about **50% of the original effect size**." Verified on the publisher version-of-record PDF via the EUR repository.
- [34] Errington, T. M., Mathur, M., Soderberg, C. K., Denis, A., Perfito, N., Iorns, E., & Nosek, B. A. (2021). *Investigating the replicability of preclinical cancer biology*. eLife 10, e71601. https://doi.org/10.7554/eLife.71601 — accessed 2026-07-29. [type: peer-reviewed] — **23 papers, 50 experiments, 158 effects.** Quoted: "the median effect size in the replications was **85% smaller** than the median effect size in the original experiments, and **92% of replication effect sizes were smaller** than the original"; overall success rate "**46% (51/112)**". Together, [31]–[34] give the deck four *differently measured* replication rates in four fields — which is the point of the slide, not the individual numbers.
- [35] GRADE Working Group — Schünemann, H., Brożek, J., Guyatt, G., & Oxman, A. (Eds.) (2013). *GRADE Handbook*. https://gdt.gradepro.org/app/handbook/handbook.html and https://www.gradeworkinggroup.org/ — accessed 2026-07-29. [type: policy] — The four certainty levels quoted verbatim on the slide, e.g. High: "We are very confident that the true effect lies close to that of the estimate of the effect"; Very low: "We have very little confidence in the effect estimate."
- [36] Balshem, H., Helfand, M., Schünemann, H. J., et al. (2011). *GRADE guidelines: 3. Rating the quality of evidence*. Journal of Clinical Epidemiology 64(4), 401–406. https://doi.org/10.1016/j.jclinepi.2010.07.015 — accessed 2026-07-29. [type: peer-reviewed] — The peer-reviewed companion to [35], and the source of the deck's key structural point: ratings apply to **bodies of evidence, not individual studies**, and **randomised trials start at high quality while observational studies start at low**. That starting-point rule is exactly what does not transfer to a field with no trials.
- [37] Sterne, J. A. C., Savović, J., Page, M. J., et al. (2019). *RoB 2: a revised tool for assessing risk of bias in randomised trials*. BMJ 366, l4898. https://doi.org/10.1136/bmj.l4898 — accessed 2026-07-29. [type: peer-reviewed] — Quoted from the Summary Points: "Bias is assessed in **five distinct domains**… These answers lead to judgments of '**low risk of bias**', '**some concerns**', or '**high risk of bias**'." Also grounds the line that the original 2008 Cochrane tool has "over 40 000 citations in Google Scholar" — i.e. this is *infrastructure*, not a preference. Verified on the version-of-record PDF via the White Rose repository; bmj.com 403s to automated fetch.
- [38] Page, M. J., McKenzie, J. E., Bossuyt, P. M., et al. (2021). *The PRISMA 2020 statement: an updated guideline for reporting systematic reviews*. BMJ 372, n71. https://doi.org/10.1136/bmj.n71 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC8005924/) — accessed 2026-07-29. [type: peer-reviewed] — **27-item checklist across seven sections**, plus an abstract checklist and revised flow diagrams. Used with [35]–[37] to make the slide's positive, verifiable claim: medicine has *named, versioned, mandated* appraisal and reporting standards. **Explicit non-claim:** the deck does **not** assert that no equivalent exists in other fields — that is an unverifiable absence claim. It asks the audience to name their own field's equivalent instead.
- [39] Patterson, D., Snyder, L., & Ullman, J. (1999). *Evaluating Computer Scientists and Engineers For Promotion and Tenure*. Computing Research Association Best Practices Memo (approved by the CRA Board, August 1999). https://cra.org/resources/best-practice-memos/evaluating-computer-scientists-and-engineers-for-promotion-and-tenure/ — accessed 2026-07-29. [type: policy] — A discipline's own professional body explaining, in writing, why its publication culture is not yours. Quoted: "The reason **conference publication is preferred to journal publication**, at least for experimentalists, is the shorter time to print (**7 months vs 1-2 years**), the opportunity to describe the work before one's peers at a public presentation, and the more complete level of review (**4-5 evaluations per paper compared to 2-3 for an archival journal**)." Age is stated on the slide; it is cited as the origin of a still-operative norm, not as current news.
- [40] arXiv (2026). *Monthly submission statistics*, and Abdill, R. J., & Blekhman, R. (2019). *Meta-Research: Tracking the popularity and outcomes of all bioRxiv preprints*. eLife 8, e45133. https://arxiv.org/stats/monthly_submissions and https://doi.org/10.7554/eLife.45133 — accessed 2026-07-29. [type: primary-doc; peer-reviewed] — The preprint-norm gap in two numbers: arXiv's own live statistics page reported **3,117,708 direct submissions** since August 1991 when checked on 2026-07-29; the peer-reviewed bioRxiv census counts "**all 37,648 preprints** uploaded to bioRxiv.org… **in its first five years**", with "**two-thirds of preprints posted before 2017** … later published in peer-reviewed journals". **Explicit non-claim:** arXiv's per-subject-area breakdown renders via JavaScript and could not be read, so the deck quotes no physics-versus-CS split.
- [41] Waltman, L. (2016). *A review of the literature on citation impact indicators*. Journal of Informetrics 10(2), 365–391. https://doi.org/10.1016/j.joi.2016.02.007 (author's accepted manuscript: https://arxiv.org/abs/1507.02099) — accessed 2026-07-29. [type: peer-reviewed] — Why a raw citation count means nothing across a field boundary, quoted verbatim: "citation counts of publications from different fields **should not be directly compared** with each other… a biochemistry publication with 25 citations cannot be considered to have a higher citation impact than a mathematics publication with ten citations. There is a difference in citation density between biochemistry and mathematics of about **an order of magnitude**." Also grounds the line that the *time constant* of citation accrual is field-dependent.
- [42] Martín-Martín, A., Thelwall, M., Orduna-Malea, E., & Delgado López-Cózar (2021). *Google Scholar, Microsoft Academic, Scopus, Dimensions, Web of Science, and OpenCitations' COCI: a multidisciplinary comparison of coverage via citations*. Scientometrics 126(1), 871–906. https://doi.org/10.1007/s11192-020-03690-4 (open preprint: https://arxiv.org/abs/2004.14329) — accessed 2026-07-29. [type: peer-reviewed] — **3,073,351 citations** to 2,515 highly-cited 2006 documents across **252 subject categories**. The finding the deck uses: coverage is not uniform across fields — Microsoft Academic "had **coverage gaps in some areas, such as Physics and some Humanities categories**", and Dimensions "displays some coverage gaps, **especially in the Humanities**". Your cross-field search is only as good as your database's coverage of the *other* field.
### D. Primary documentation — the cross-field discovery, mapping and shared-knowledge tools
*Every cell of the Cross-Field Discovery Toolkit table on slides 18–19 and in the handout comes from the pages below, fetched 2026-07-29. Nothing in that table is taken from a review, a listicle or a vendor blog. Where a vendor does not document something, the table says "not documented".*
- [43] Connected Papers (2026). *About* and *Pricing*. https://www.connectedpapers.com/about and https://www.connectedpapers.com/pricing — accessed 2026-07-29. [type: primary-doc] — The deck's single most useful piece of tool documentation, because it tells you what the picture is *not*: "In the graph, papers are arranged according to their similarity… even papers that do not directly cite each other can be strongly connected and very closely positioned. **Connected Papers is not a citation tree.**" Method quoted: "Our similarity metric is based on the concepts of **Co-citation and Bibliographic Coupling**… two papers that have highly overlapping citations and references are presumed to have a higher chance of treating a related subject matter." Scale: "To create each graph, we analyze an order of **~50,000 papers** and select the few dozen with the strongest connections to the origin paper." Corpus: "connected to the **Semantic Scholar Paper Corpus** (licensed under ODC-BY)". Access as of 2026-07-29: Free tier "**5 graphs per month**", "All features included"; Academic "**USD3/month**", billed 36 USD annually, unlimited graphs; Business USD10/month. **Not documented:** exact corpus size (only "hundreds of millions"), snapshot update frequency, group seat prices.
- [44] Inciteful (2026). *Inciteful — Paper Discovery, Literature Connector and Data Sources*. https://incitefulmed.com/academic/ , https://incitefulmed.com/academic/c and https://incitefulmed.com/academic/data — accessed 2026-07-29. [type: primary-doc] — **The one tool in the table whose own documentation names interdisciplinary work as the use case**, which is why it anchors rung 2 of the Ladder. Quoted: "**Interested in interdisciplinary studies?** Discover how two bodies of literature connect to one another through citations… Enter two papers and we'll show you the **shortest paths** between them." Mechanics quoted: "A link is established through citations. Two papers are linked if either one cites the other… Starting from the seed paper, we recursively search through all the papers that are cited by the seed paper or that cite the seed paper"; the hop limit is "six, but we've honestly not found two papers that are more than five away". Method quoted: "we build a citation network centered around your paper(s)… you find not only the most 'important' papers in the graph, but also the most similar using **link prediction algorithms** — the same algorithms used in social networks to suggest friends." Scale and cost, verbatim tiles: "**240M+** Academic papers", "**2B+** Citations indexed", "**100% Free to use**". Corpus: **OpenAlex, Semantic Scholar, Crossref and OpenCitations**. **Domain note stated in the handout:** `inciteful.xyz` now 301-redirects to `incitefulmed.com/academic/`. **Not documented:** non-shortest paths ("Not yet") and multi-paper endpoints ("Not right now") — both stated by the FAQ as absent.
- [45] ResearchRabbit (2026). *Home*, *Pricing* and *Help guide*. https://www.researchrabbit.ai/ , https://www.researchrabbit.ai/pricing and https://www.researchrabbit.ai/help/guide — accessed 2026-07-29. [type: primary-doc] — Carried forward from Session 2's Literature-Discovery Tool Matrix and re-checked. Corpus: "Access over **310 million academic papers**". Discovery modes quoted: "**Similar Work**", "**Earlier Work**" ("the references of specific papers"), later work, "These Authors" and "Suggested Authors"; per-paper controls "Similar, References, and Cited By. Each one opens a new map". Graph semantics quoted: "X-axis = timeline (older to newer papers)", "Y-axis = influence (citation count)". **Not a correction to Session 2 — a promotion from its handout to a slide.** Session 2's handout already recorded the cap ("Free Forever tier (up to 50 seed articles)", `session-2-literature/handout.md`); it simply never reached Session 2's slide, and attendees reliably miss handout-only detail. Re-verified here: the free tier is real ("**$0, Forever!**", including "Collaborate by sharing your collection") but capped at "**Use up to 50 seed articles**"; RR+ at **$10/month annual or $12.50 monthly** raises that to 300. **Not documented:** which upstream provider supplies the 310M records; whether shared collections are co-editable.
- [46] Litmaps (2026). *Pricing* and *Documentation*. https://www.litmaps.com/pricing and https://docs.litmaps.com/ — accessed 2026-07-29. [type: primary-doc] — Corpus quoted: "The Litmaps database has **270+ million research articles**", ingested from "multiple data providers" named as **Crossref, Semantic Scholar and OpenAlex**, and "**updates the database weekly**". Free tier: "**2 Litmaps**", "**100 articles per Map**", "Basic Search Up to 20 inputs". Pro: "**$10*/month**", billed annually at $120/year, unlimited maps and articles. Seed-map behaviour quoted: the seed map finds papers that "either cite or are cited by" the seed, and Discover "accommodates multiple seed inputs". **Not documented:** the Team-tier price; the ranking algorithm behind Discover beyond "connection, using citations and references".
- [47] Open Knowledge Maps (2026). *About* and *FAQ*. https://openknowledgemaps.org/about and https://openknowledgemaps.org/faq — accessed 2026-07-29. [type: primary-doc] — The free, non-profit option in the table, and the only one that maps a *query* rather than a *seed paper*. Corpora quoted: "either the **PubMed API or the BASE API**". Map size quoted: "We want to keep the number of resources to a manageable amount. **100 resources** are already 10 times more content than is presented on a standard search results page." Method quoted: the system analyses "titles, abstracts, authors, journals, and subject keywords to create a **word co-occurrence matrix** between articles. On top of this matrix, we perform clustering and ordination algorithms." Software is open source; the organisation is "a charitable non-profit organization".
- [48] VOSviewer (2026). *Download* and *Features*. https://www.vosviewer.com/download and https://www.vosviewer.com/features/highlights — accessed 2026-07-29. [type: primary-doc] — "a software tool for constructing and visualizing bibliometric networks", building co-authorship, citation, **bibliographic coupling**, **co-citation** and **term co-occurrence** networks, with data from Web of Science, Scopus, Dimensions, Lens, PubMed and API retrieval from Crossref, Europe PMC and OpenAlex. Current release quoted: "**VOSviewer version 1.6.21, released on June 12, 2026**". **Licence caveat stated in the handout:** the official wording is only "The software **can be used freely for any purpose**" — the deck therefore says "free to use", never "open source".
- [49] OpenAlex (2026). *Topics*, *Semantic search*, *Authors* and *Authentication & pricing* (developer documentation), plus the live API. https://developers.openalex.org/api-reference/topics , https://developers.openalex.org/ and https://api.openalex.org/ — accessed 2026-07-29. [type: primary-doc] — The classification backbone of the discovery section. Hierarchy quoted: "Topics… exist in a four-level hierarchy: **domain > field > subfield > topic**", with a documented table of 4 / 26 / 254 / ~4,500. **Independently re-counted against the live API on 2026-07-29** by reading `meta.count` from `/domains`, `/fields`, `/subfields` and `/topics`: **4 domains, 26 fields, 252 subfields, 4,516 topics.** The docs/API mismatch on subfields (254 vs 252) is stated on the slide rather than papered over. Cross-field navigation mechanism quoted: "Every work is assigned a `primary_topic` which includes the full hierarchy path. You can filter by any level: `filter=primary_topic.domain.id:1`…". Semantic search quoted: "OpenAlex embeds the title and abstract of every work using **GTE Large EN**… into a **1,024-dimensional vector**. At query time we embed your query the same way and return the works closest by **cosine similarity**", with a worked cross-field example in the docs — a query about "predicting drug toxicity from molecular structure" finding papers that say "computational toxicology" or "QSAR", "words your search never mentioned". Limits quoted: 2,000-character input, 50 results per query, 1 request per second. **Access change stated on the slide, because it is new and affects planning:** OpenAlex now runs usage-based pricing — "Your free API key gives you **$1 of free usage every day**", $0.10/day without a key, with semantic search billed at "$1" per 1,000 calls, and "the openalex.org website runs on this same API, so **browsing it draws from the same budget**". **Not documented:** the prices of the Premium/Institutional/Partner tiers (openalex.org/pricing returns 403), and the topic-classification model itself — the docs say only that topics are "automatically assigned", so the deck says nothing further about how.
- [50] Semantic Scholar (2026). *Academic Graph API* and *Recommendations API* documentation. https://api.semanticscholar.org/api-docs/graph and https://api.semanticscholar.org/api-docs/recommendations — accessed 2026-07-29. [type: primary-doc] — Carried forward from Session 2 and extended with the two endpoints that matter for cross-field work. Recommendations quoted: `POST /papers/` "Get recommended papers for **lists of positive and negative example papers**", taking `positivePaperIds` and `negativePaperIds` — the documented way to push results *away* from your home field. `GET /papers/forpaper/{paper_id}` takes a `from` parameter, "Which pool of papers to recommend from", with the enum "recent" or "all-cs"; `limit` default 100, "Maximum 500". SPECTER2 access quoted from the Graph docs: "Specify `embedding.specter_v2` to select v2 embeddings." Corpus figures on the product page: 214 million papers, 2.49 billion citations. The endpoint's live behaviour was confirmed by the author on 2026-07-29 with an unauthenticated call. **Explicit non-claim:** the Recommendations spec does **not** state that it uses SPECTER2, so the deck does not say it does.
- [51] Elsevier (2024/2026). *What is the complete list of ASJC subject areas in Scopus?* and *What are the Scopus subject area categories?* (Scopus Support Center). https://service.elsevier.com/app/answers/detail/a_id/15181/supporthub/scopus/ — accessed 2026-07-29. [type: primary-doc] — Grounds the "who decides what field a paper is in" slide. Quoted: "Serial titles are classified using the ASJC (All Science Journal Classification) scheme. This is done by **in-house experts at the moment the serial title is set up** for Scopus coverage; the classification is based on the aims and scope of the title, and on the content it publishes." **Counting note stated on the slide:** Elsevier's page never states a total; the figures used (334 four-digit codes, 27 top-level subject areas) were **counted from Elsevier's own published table on 2026-07-29** and are presented as such, not as an Elsevier statement.
- [52] Clarivate (2025). *Web of Science Core Collection: Web of Science Categories*. https://support.clarivate.com/ScientificandAcademicResearch/s/article/Web-of-Science-Core-Collection-Web-of-Science-Categories — accessed 2026-07-29. [type: primary-doc] — The other half of the same slide, quoted verbatim: "**Web of Science Categories are assigned at the journal level.** All items in a journal will be assigned the Web of Science Categories of the journal it is published in"; "Every journal covered by Web of Science Core Collection is assigned one or more Web of Science Categories. **A journal may have up to 6 categories** assigned to it." **Explicit non-claim:** the total number of categories is *not* stated on any first-party page that would load (the linked help file now 403s behind a challenge), so the deck gives the assignment rule and says the count is not documented.
- [53] Zotero (2026). *Groups* and *Storage*. https://www.zotero.org/support/groups and https://www.zotero.org/storage — accessed 2026-07-29. [type: primary-doc] — The shared-knowledge-base row of the team-science slide. Three group types quoted (Private; Public, Closed Membership; Public, Open Membership). Quoted: "**There is no limit on how many members may join your groups**"; and the fact that decides who pays — "Group file storage **always draws from the storage account of the group owner**", so "Other group members' storage quotas will not be affected by the group files." Tiers as of 2026-07-29: 300 MB free; 2 GB $20/year; 6 GB $60/year; unlimited $120/year.
- [54] Google (2026). *Gemini Notebook (formerly NotebookLM) — FAQ* and *Upgrade your plan*. https://support.google.com/gemininotebook/answer/16269187 and https://support.google.com/gemininotebook/answer/16213268 — accessed 2026-07-29. [type: primary-doc] — Carried forward from Session 3 and re-checked; the help centre is now titled **Gemini Notebook**, so the deck names both. Free-tier limits quoted: "Get **100 notebooks, with up to 50 sources each** and 500,000 words each"; "**daily limits of 50 chat queries** and 3 audio generations"; "The current limit is 500,000 words per source or up to **200MB** for local uploads." Sharing: Viewer versus Editor roles, with a lock/shared/globe indicator; "Your data will never be shared by Gemini Notebook." **Not documented on these pages:** the dollar prices of the Plus/Pro/Ultra tiers. **Re-verified 2026-08-21** (Google Workspace Updates blog + the same support pages): the rename to Gemini Notebook was announced 16 July 2026 and the free-tier limits above are unchanged (100 notebooks / 50 sources / 500,000 words / 50 chat queries and 3 audio generations per day).
- [55] NIH (2026). *NIH RePORTER APIs*. https://api.reporter.nih.gov/ — accessed 2026-07-29. [type: primary-doc] — The collaborator-discovery row that is not a citation database. Quoted: the APIs "are designed to programmatically expose relevant scientific awards data from both NIH and non-NIH federal agencies"; the database "is available to all public users". Rate guidance quoted: "no more than **one URL request per second**", max 500 records per request. Used on the slide for one point only: funded-project records surface who is *currently working* in an adjacent field, which is earlier signal than publications.
- [56] ORCID (2026). *What is ORCID?*. https://info.orcid.org/what-is-orcid/ — accessed 2026-07-29. [type: primary-doc] — Quoted: ORCID is "a global, **not-for-profit** organization" and the iD is "a unique, persistent identifier **free of charge** to researchers", with a Public API. Used only to ground the claim that a cross-field author search has a free, non-commercial identifier layer under it. **Not documented on this page:** the number of registered iDs.
### E. The demo's field pair — clinical prediction research ↔ machine learning
*The demo corpus is eight real papers whose DOIs were verified live in Crossref on 2026-07-29; seven are open access. The seven entries below are the ones a slide or the handout quotes directly; the full eight-paper list, with access status per paper, is in `demo-script.md` and `handout.md`.*
- [57] Collins, G. S., Moons, K. G. M., Dhiman, P., et al. (2024). *TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods*. BMJ 385, e078378. https://doi.org/10.1136/bmj-2023-078378 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC11019967/) — accessed 2026-07-29. [type: peer-reviewed] — Reporting guideline; the source of three cells of the false-friend table and the demo's rung-1 check. Definitions quoted verbatim: discrimination is "How well the predictions from the model differentiate between individuals with and without the outcome"; calibration is "Agreement between observed outcomes and estimated values from the model"; internal validation is "Evaluating the performance of a prediction model on the same population on which the model was developed (eg, train test split, cross validation, or bootstrapping)". **The deck's strongest single piece of false-friend evidence** is the guideline's own decision to abandon the contested word: "we refer to **validation as evaluation** in this article".
- [58] Van Calster, B., McLernon, D. J., van Smeden, M., Wynants, L., & Steyerberg, E. W. (2019). *Calibration: the Achilles heel of predictive analytics*. BMC Medicine 17, 230. https://doi.org/10.1186/s12916-019-1466-7 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC6912996/) — accessed 2026-07-29. [type: peer-reviewed] — Corroborates [57] on both terms in a peer-reviewed research article rather than a guideline. Quoted: "Algorithms (or risk prediction models) should give **higher risk estimates for patients with the event than for patients without the event ('discrimination')**. Typically, discrimination is quantified using the area under the receiver operating characteristic curve (AUROC or AUC), also known as the concordance statistic or c-statistic"; and "The accuracy of risk estimates, relating to the agreement between the estimated and observed number of events, is called '**calibration**'."
- [59] Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). *Dissecting racial bias in an algorithm used to manage the health of populations*. Science 366(6464), 447–453. https://doi.org/10.1126/science.aax2342 — accessed 2026-07-29. [type: peer-reviewed] — Grounds the *other* meaning of both "discrimination" and "bias", in a paper about a deployed model with good predictive accuracy. Quoted: "a widely used algorithm… exhibits significant **racial bias**: At a given risk score, Black patients are considerably sicker than White patients… Remedying this disparity would increase the percentage of Black patients receiving additional help **from 17.7 to 46.5%**"; and the mechanism, "the algorithm predicts health care **costs** rather than **illness**". The demo uses it as the worked example of a model that is excellent on field A's metric and unacceptable on field B's. Abstract verified via Europe PMC (PMID 31649194); science.org 403s to automated fetch.
- [60] Christodoulou, E., Ma, J., Collins, G. S., Steyerberg, E. W., Verbakel, J. Y., & Van Calster, B. (2019). *A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models*. Journal of Clinical Epidemiology 110, 12–22. https://doi.org/10.1016/j.jclinepi.2019.02.004 — accessed 2026-07-29. [type: peer-reviewed] — **71 of 927 screened studies, 282 model comparisons.** Grounds the false-friend table's "calibration is a missing concept" cell, quoted: "**Calibration was not addressed in 56 (79%) studies**"; and "In 48 (68%) studies, we observed **potential bias in the validation procedures**". Headline conclusion quoted in the demo: "We found **no evidence of superior performance of ML over LR**", with the difference in logit(AUC) at low risk of bias being "0.00 (95% confidence interval, −0.18 to 0.18)". Abstract verified via Europe PMC (PMID 30763612).
- [61] Assel, M., & Vickers, A. (2025). *The F score ranks diagnostic tests and prediction models inconsistently with their clinical utility*. Diagnostic and Prognostic Research 9(1), 30. https://doi.org/10.1186/s41512-025-00214-7 — accessed 2026-07-29. [type: peer-reviewed] — The two synonym rows of the false-friend table, from a peer-reviewed source that frames the split as a *disciplinary* one rather than a coincidence. Quoted: "A similar trend is seen in **discipline specific language**, with the terms 'precision' and 'recall' — **what are more typically described in medical contexts as positive predictive value and sensitivity** — being far more commonly referenced in the contemporary literature"; and "Fβ is calculated from **precision (more commonly termed positive predictive value)** … and **recall (sensitivity)**". Preferred over a software library's documentation, which was the earlier candidate and is now cited nowhere in this deck.
- [62] Paulus, J. K., & Kent, D. M. (2020). *Predictably unequal: understanding and addressing concerns that algorithmic clinical prediction may increase health disparities*. npj Digital Medicine 3, 99. https://doi.org/10.1038/s41746-020-0304-9 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC7393367/) — accessed 2026-07-29. [type: peer-reviewed] — A peer-reviewed paper that names the false friend explicitly, which is why the deck can assert the clash rather than merely juxtapose two definitions. Quoted verbatim: "**Disentangling the two meanings of 'discrimination' — discernment between individuals' risk of a future event on the one hand and unfair prejudice leading to inequity on the other** (akin to what economist Thomas Sowell has referred to as Discrimination I and Discrimination II, respectively) — is central to understanding algorithmic fairness, and more deeply problematic than generally appreciated."
- [63] Sung, J., & Hopper, J. L. (2023). *Co-evolution of epidemiology and artificial intelligence: challenges and opportunities*. International Journal of Epidemiology 52(4), 969–973. https://doi.org/10.1093/ije/dyad089 (open: https://pmc.ncbi.nlm.nih.gov/articles/PMC10396412/) — accessed 2026-07-29. [type: peer-reviewed] — The single best false-friend source found for any field pair, and the model for what a rung-1 glossary should look like. Quoted on the "bias" row: "In epidemiology, bias means **systematic errors in statistical inference which arise from weaknesses in study designs and conduct**… **In AI, bias means the intercepts in unit models**" — a three-way clash in one sentence, before the fairness sense of [59] and [62] is even added. Its Table 1, *Comparative terminology between epidemiology and artificial intelligence for similar concepts described by different terms*, supplies the deck's synonym mappings: **internal validation ↔ validation**, **external validation ↔ test**, risk factors ↔ features, model fitting ↔ model training, winner's curse ↔ overfitting. Also quoted in the handout: "The 'Dictionary of Epidemiology' **lacks the term 'fairness'**."
### F. Added in the August 2026 revision — current-generation companion evidence
*Added 2026-08-21 to pair the deck's labelled 2024–25 baselines with current-generation evidence. Accessed dates are the actual fetch dates of the revision pass.*
- [64] Peters, U., Bertazzoli, A., DeJesus, J. M., van der Velden, G. J., & Chin-Yee, B. (2026). *Generics in science communication: Misaligned interpretations across laypeople, scientists, and large language models*. Public Understanding of Science, first published online 20 April 2026. https://doi.org/10.1177/09636625261425891 — accessed 2026-08-21. [type: peer-reviewed] — The same team's direct follow-up to [22], one model generation newer (**ChatGPT-5, DeepSeek-V3.1**). Reported effects: ChatGPT-5 rated generics as **more** generalizable than laypeople (b = .25, p = .009), DeepSeek-V3.1 more so (b = .53, p < .001); human domain experts went the *other* way (psychologists b = −.32, biomedical researchers b = −.48, both p < .001). The teaching point used on slide 12: experts read a generic claim *narrowly*; LLMs read it *more broadly than even untrained laypeople* — the mechanism behind generalization bias, confirmed on a GPT-5-class model. Still one generation behind the August-2026 cohort, and labelled as such.
- [65] Artificial Analysis (2026). *MMLU-Pro* and *Humanity's Last Exam* evaluation pages (independent benchmark runs; the HLE run covers 2,158 text-only questions). https://artificialanalysis.ai/evaluations/mmlu-pro and https://artificialanalysis.ai/evaluations/humanitys-last-exam — accessed 2026-08-21. [type: primary-doc] — Grounds the two current-generation numbers on slide 13, both attributed to Artificial Analysis on the slide: frontier MMLU-Pro scores cluster at **83–90%** (top listed: Gemini 3 Pro Preview 89.8%, Claude Opus 4.5 (Reasoning) 89.5%) — near-saturation, which is why the deck retires [18]'s 2024 spread to a labelled baseline; and the best tracked HLE score, **Claude Fable 5 at 55.5%**. An independent evaluator, not vendor self-reporting — but a leaderboard, so both figures carry a date on the slide and should be re-checked before each delivery.
- [66] Onweller, H., Lumer, E., Huber, A., Ramchandani, P., Subbiah, V. K., & Feld, C. (2026). *Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents*. **Preprint** (submitted 7 May 2026). https://arxiv.org/abs/2605.06635 — accessed 2026-08-21. [type: peer-reviewed] — Labelled a preprint on the slide. **14 closed- and open-source current-generation models** (incl. Claude Opus 4.5/4.6, GPT-5.2/5.4, Gemini 3.1 Pro) audited for link accessibility, topical alignment and factual accuracy. Quoted: "even the strongest frontier models maintain **link validity above 94%** and relevance above 80%" yet achieve only "**39–77% factual accuracy**"; "Providers that generate more citations achieve lower factual accuracy." The current-generation companion to [21]: the dominant failure has moved from *fake URL* to *real URL, unsupported claim* — the evidential basis for promoting rung 3's read-two-in-full check over the resolve check, on slides 13 and 25 and in the demo's §5a.
---
**Counts for this session:** 66 annotated sources — **46 peer-reviewed/preprint** ([1]–[14], [16]–[29], [31]–[34], [36]–[38], [41], [42], [57]–[64], [66]), **4 policy** ([15], [30], [35], [39]), **16 primary-doc** ([40], which pairs arXiv's own statistics page with a peer-reviewed bioRxiv census, plus [43]–[56] and [65]). Six are explicitly labelled **preprints** on the slide that uses them ([20], [25], [26], [27], [28], [66]). That is well past the series minimum of 15 sources and 5 peer-reviewed. Every statistic on a slide resolves to an entry above; every DOI in sections A–E was verified live against Crossref on 2026-07-29, and every URL in those sections was opened on the same day. Section F's pages were fetched on 2026-08-21 during the August revision; sources kept unchanged keep their original accessed dates.
**Independently verified by the author, not taken from a secondary source.**
- The OpenAlex hierarchy counts on slides 16 and 18 (**4 domains, 26 fields, 252 subfields, 4,516 topics**) were read from `meta.count` on the live `/domains`, `/fields`, `/subfields` and `/topics` endpoints on 2026-07-29. OpenAlex's own documentation table says 254 subfields; the mismatch is stated on the slide rather than hidden.
- The Semantic Scholar Recommendations endpoint was called unauthenticated on 2026-07-29 and returned results, confirming the "no key needed" claim in [50].
- The per-discipline spread on slide 13 (**Psychology 79.2% → Law 51.0% for GPT-4o**) was read directly from Table 2 of the MMLU-Pro paper's arXiv HTML, not from a summary of it. The current-generation companions on that slide (MMLU-Pro frontier cluster **83–90%**; HLE **55.5%**, Claude Fable 5) were read from Artificial Analysis's evaluation pages on 2026-08-21 [65].
- The ASJC figures under [51] (334 four-digit codes, 27 top-level areas) were **counted from Elsevier's own published table**; Elsevier states no total, and the slide says so.
- All eight demo-corpus DOIs, plus every DOI in sections A–C, were checked against the Crossref REST API on 2026-07-29 and returned matching titles, venues and volumes.
**Dropped during research, and why.**
(a) **Baker, M. (2016), "1,500 scientists lift the lid on reproducibility", Nature 533:452–454.** The article is real and the metadata was verified, but its survey percentages — the widely quoted 52% / 70% / by-discipline breakdown — sit behind Nature's paywall and could not be read on any primary page. It is also a self-selected online poll, not a probability sample. **Dropped entirely**; [31]–[34] carry the replication argument in peer-reviewed form instead.
(b) **A "no GRADE-equivalent exists outside medicine" claim.** This is an unverifiable absence claim. The deck states positively what medicine has standardised ([35]–[38]) and asks the audience to name their own field's equivalent.
(c) **Bromham et al.'s "X% lower funding success".** No percentage-point figure appears in the abstract, so none appears in this deck.
(d) **Wang, Thijs & Glänzel per-standard-deviation effect sizes** (1.48% / 2.45% / 5.77%) — surfaced during research but not traceable to the results tables. Not used.
(e) **Star & Griesemer (1989) on boundary objects.** Bibliographically real and highly relevant, but the abstract could not be read on any loadable page (SAGE 403s; the open scan is unparseable). Rather than quote it on trust, the deck uses [11] for the same argument.
(f) **SPECTER2's frequently repeated "6M triplets across 23 fields of study."** Traceable only to a vendor blog post. Not used; [50] documents only what the API spec states.
(g) **OpenAlex's topic-classification model** (an mBERT variant, with top-1 accuracy figures). Widely repeated, not documented on any first-party page reachable on 2026-07-29. The slide says "not documented".
(h) **"Web of Science has 252 subject categories."** Clarivate's own article states the assignment *rule* but no total, and the linked category list now 403s. The slide gives the rule and says the count is not documented.
(i) **Elicit/SciSpace/Consensus effectiveness figures from a 2024 LIS journal article** — the article carries a retraction notice in the journal's own listing. Excluded.
(j) **FIRE-Bench** (an agentic rediscovery benchmark). Real and verified, but all 30 of its tasks are inside machine-learning research, so despite its framing it says nothing about *cross-domain* performance. Out of scope.
(k) **arXiv's per-subject-area submission breakdown.** The official page renders its tables in JavaScript and could not be read, so the deck quotes only the single total from [40].
(l) **scikit-learn's documentation as the source for "recall = sensitivity" and "precision = PPV".** It was fetched and it does say both, but a peer-reviewed source that frames the split as *disciplinary* is strictly better for this audience, so [61] replaced it and the library is cited nowhere in this deck.
(m) **Bzdok, Altman & Krzywinski (2018), *Statistics versus machine learning*, Nature Methods 15:233–234.** Real, open and excellent on the inference-versus-prediction split, but a full-text search found no terminology or glossary content at all — so it is **not** used as a false-friend source. [63] carries that argument instead. It stays in the handout's further reading, where its actual contribution belongs.
(n) **Rajkomar, Dean & Kohane (2019), *Machine Learning in Medicine*, NEJM 380:1347–1358.** Verified and apt, but fully closed access, so it is not in an open reading list; the 2025 BMJ Oncology review replaces it.
(o) **Notion, Scopus author profiles, and medRxiv's own statistics** — considered for the team-science and collaborator rows, but their official pages either were not reached or returned 403 on 2026-07-29. Omitted rather than described from memory.
---
*AI for Researchers · Session 6: AI for Multidisciplinary Research · Landscape as of August 2026*