AI for Researchers

Session 6: AI for Multidisciplinary Research

Getting oriented in a field that is not yours — in four rungs, with the checks that keep it honest

{{Presenter Name}}

Landscape as of August 2026

2-Minute Recap

The Series So Far


Today (Session 6): the job no colleague has time to do for you — getting oriented in someone else's field, fast enough to be useful and checked enough to be safe.

Learning Objectives

By the End of This Session You Will Be Able To…


  1. Apply AI tools to translate an unfamiliar discipline's jargon, methods, and core literature into accessible terms.
  2. Run a citation-graph or cross-field semantic search to find adjacent literature and unexpected connections.
  3. Evaluate evidence from fields with different methodological standards for compatibility before synthesizing it.

Objectives 4–6 continue on the next slide.

Learning Objectives

…And You Will Also Be Able To


  1. Apply AI to build a cross-disciplinary synthesis memo combining evidence from multiple fields.
  2. Explain how AI can support team science: shared knowledge bases, collaborator discovery, and cross-expertise communication.
  3. Produce an annotated reading list to orient into an unfamiliar field.

Section 01

Why Crossing Fields Is Hard

Not a confidence problem, and not a you problem. A measured, structural one — and knowing its exact shape tells you what to delegate.

Literacy Foundation — Start With the Prize

The Highest-Impact Science Is Mostly Conventional, With One Foot Outside


"Papers of this type were twice as likely to be highly cited works. Novel combinations of prior work are rare, yet teams are 37.7% more likely than solo authors to insert novel combinations into familiar knowledge domains." Uzzi, Mukherjee, Stringer & Jones, Science 2013 — 17.9 million papers, all fields [1]

The Evidence — Measured, Not Asserted

And It Is Priced In: Four Bibliometric Penalties


18,476grant proposals: "the greater the degree of interdisciplinarity, the lower the probability of being funded" [6]
8 vs 20+years to half of a cohort ceasing to publish — most interdisciplinary vs moderately so, 154,021 biomedical PhDs [7]
128,950manuscripts incl. rejections: interdisciplinary topics lowered acceptance; interdisciplinary references raised it [8]

Read with the caveat its own field states: interdisciplinarity indicators give "surprisingly deviant results" and "should be interpreted with much caution". [13] The direction is consistent across studies; the effect sizes are not. Practical reading of [8]: cite across fields, frame within one.

Critical Literacy — The Key Slide

The Mechanism Is Language. Then, Underneath It, Epistemology.


"The unique languages of the disciplines reflect deeper differences in underlying assumptions, epistemologies (ways of knowing), philosophies, and approaches to science and societal problems." National Research Council, Enhancing the Effectiveness of Team Science, 2015 [11]

That last number is about humans, not models. Thirty minutes of searching does not make you a non-expert who can read field B — it makes you a non-expert with tabs open. This session is about the thirty minutes that does work.

Section 02

AI as Field-Translator

Rung 1 of today's workflow — the one thing AI is unusually good at, and the one place it will mislead you most convincingly.

Literacy Foundation

"Translate This Field" Is Three Different Requests


A translation you cannot audit is not a translation. It is a feeling of comprehension — and the next two slides are about how reliably AI produces that feeling without the comprehension.

The Terminology Trap — Worked on One Real Field Pair

False Friends: Words That Cross the Border and Change Meaning


The word In clinical prediction research In machine learning Why it bites
Discrimination Desirable. "How well the predictions… differentiate between individuals with and without the outcome" [57] [58] Undesirable. "Unfair prejudice leading to inequity" [62] [59] A model can have excellent discrimination and be discriminatory — both true at once
Bias "Systematic errors… which arise from weaknesses in study designs and conduct" [63] "In AI, bias means the intercepts in unit models" [63] — plus the fairness sense [59] Three meanings, one word. "Low bias" is praise in one, and silent on the other two
Validation External validation — so contested that TRIPOD+AI "refer[s] to validation as evaluation" throughout [57] "data used for parameter tuning" [57]; epi's external validation is ML's test set [63] "We validated the model" names two very different amounts of evidence
Calibration "Agreement between observed outcomes and estimated values" — routinely required [57] [58] Often absent: not addressed in 56 of 71 (79%) ML-vs-regression studies [60] Not a different meaning — a missing concept. The hardest kind of gap to notice
Sensitivity= recall Sensitivity: of those with the outcome, the share the model flags Recall — one of the terms "more typically described in medical contexts as positive predictive value and sensitivity" [61] Same quantity, two names. Search one, find nothing
PPV= precision Positive predictive value: of those flagged, the share that have it Precision — "more commonly termed positive predictive value" [61] The pair that derails a methods conversation fastest
One field pair, six words, every cell quoted from a reporting guideline or a peer-reviewed paper — never from a model. Build this table for your own field pair first: it is rung 1 of the Ladder, and it is the deliverable.

The Evidence — Two Results That Should Change How You Prompt

It Reads Better Than It Informs


"While LLMs can generate [plain-language summaries] that appear indistinguishable from human-written ones in subjective evaluations, human-written [ones] lead to significantly better comprehension." Guo, Sohn, Leroy & Cohen, J. Biomedical Informatics 2026 — 150 participants, pre-2026 models, comprehension and recall tested [23]

So the check on rung 1 is not "does this read clearly?" — clarity is the failure mode. It is: can I find each of these twenty definitions in one field-B review article?

Critical Literacy — Session 1's Rule 2, Quantified

It Is Weakest Exactly Where You Are Weakest


6% → 29%fabricated citations, well-covered vs niche topic — same model, same task (2025 study) [21]
79% → 51%GPT-4o, 2024 — best discipline to worst on a benchmark today's frontier near-saturates at 83–90% [18] [65]
0.64best Macro-F1 at simply identifying interdisciplinary papers (random = 0.18; 2025 cohort) [26]

The live per-discipline evidence is now Humanity's Last Exam: "low accuracy and calibration" at the expert frontier [19]; best tracked score 55.5% (Claude Fable 5, Artificial Analysis, Aug 2026) [65] — better, still far from expert level. [20], [25] and [26] are preprints on pre-2026 cohorts. It is all Session 1's Rule 2: rarity moves tasks rightward.

CRIT, Applied — Rung 1

The Translation Prompt, Built to Be Audited


"No papers, no web search" is deliberate. Rung 1 is a vocabulary task; with retrieval on by default, the model will otherwise reach for pages you cannot audit yet — and citations are Session 1's Dangerous zone: [21] is why. Papers arrive on rung 2, from a database, or not at all.

Section 03

Finding the Adjacent Literature

Rung 2 — where AI stops being the source and becomes the index. Three search geometries, and one question about who decided what field your paper is in.

Literacy Foundation

Three Search Geometries — and Who Decides What Field a Paper Is In


So an interdisciplinary paper in a disciplinary journal inherits its venue's fields — invisible to a field-B browse. Coverage differs too: gaps cluster "especially in the Humanities". [42]

Series Artefact — Extending Session 2's Matrix

The Cross-Field Discovery Toolkit (1 of 2): Map the Landscape


Tool What it maps, and from what Access, at last check A limit it documents itself
Connected Papersone seed paper A similarity graph from co-citation + bibliographic coupling; "~50,000 papers" analysed per graph [43] Free: 5 graphs/month. Academic USD3/mo [43] "Connected Papers is not a citation tree" — adjacency is similarity, not citation [43]
IncitefulLiterature Connector Two seeds: "how two bodies of literature connect… through citations". 240M+ papers [44] "100% Free to use"; on OpenAlex, S2, Crossref, OpenCitations [44] Shortest paths only; one paper per end; max 6 hops [44]
ResearchRabbitcarried from Session 2 Similar / Earlier / Later work and author networks; 310M+ papers [45] "$0, Forever!", capped at 50 seeds. RR+ $10/mo → 300 [45] Upstream corpus provider not documented [45]
Litmapsmulti-seed Papers that "either cite or are cited by" the seed; 270M+ articles [46] Free: 2 maps, 100 articles. Pro USD10/mo [46] Ranking behind "Discover" not documented [46]
Open Knowledge Mapsno seed paper needed A query, clustered from a word co-occurrence matrix over titles, abstracts, keywords [47] Free; non-profit; open source. PubMed or BASE [47] 100 documents per map, by design — "a manageable amount" [47]
VOSviewerbring your own export Co-citation, coupling and term co-occurrence networks from a WoS / Scopus / PubMed export [48] "can be used freely for any purpose"; v1.6.21, June 2026 [48] "Free to use" — the site names no open-source licence, so nor does this deck [48]
Landscape as of August 2026; every cell from that tool's own documentation, fetched 2026-07-29. Where a vendor documents nothing, this table says so. Not a ranking.

Series Artefact — Extending Session 2's Matrix

The Cross-Field Discovery Toolkit (2 of 2): Search Sideways, Find People


Capability What it does, documented Access, at last check A limit it documents itself
OpenAlex semantic searchpaste a paragraph, not a phrase Embeds every title + abstract into a 1,024-dim vector: "the richer the input, the better the matches" [49] Usage-priced since 2026: "$1 of free usage every day" with a key [49] 2,000-char input, 50 results, 1 req/s — and browsing the website spends the same budget [49]
Semantic Scholar Recommendationsthe "unlike this" lever Takes positive and negative example papers — the documented way to push results away from your field [50] Free; most endpoints need no key. 214M papers [50] Pool is only "recent" or "all-cs"; the spec does not say it uses SPECTER2 [50]
OpenAlex topic filtersmove one level up Every work carries a primary_topic with its full path: 4 domains / 26 fields / ~252 subfields / ~4,500 topics [49] Free key tier as above; List+Filter is the cheapest class [49] The classification model is not documented — the docs say only "automatically assigned" [49]
OpenAlex Authors · ORCIDwho works in field B Filter author profiles by topic, institution, country or ORCID [49] [56] ORCID: "free of charge to researchers" [56] Author queries are billable List+Filter calls [49] [56]
NIH RePORTERwho is funded right now Awards "from both NIH and non-NIH federal agencies" — earlier signal than papers [55] "available to all public users"; free; bulk download [55] US federal awards only; "no more than one URL request per second" [55]
Landscape as of August 2026, from first-party documentation fetched 2026-07-29 unless noted. Two recent changes catch people out: OpenAlex moved to usage-based pricing, and Google renamed NotebookLM to Gemini Notebook (16 July 2026; rename and free-tier limits re-verified 2026-08-21).

Critical Literacy

Why Rung 2 Requires Two Tools


[27] and [28] are preprints; [16] is peer-reviewed in Nature. Read together they are the argument for this whole session: AI makes you faster at crossing fields, and makes science narrower — unless you use it to reach the people and papers it would otherwise hide.

Section 04

Synthesis Across Evidence Standards

Rungs 3 and 4 — where "significant", "validated" and "replicated" turn out to name seven different things, and combining them naively is the failure mode.

The Evidence

"A Significant Result" Names at Least Seven Different Bars


Field The threshold convention A published replication figure Where the work appears
Particle physics 5-sigma — "roughly a P value threshold of 3 × 10⁻⁷" [29] Journals and arXiv: 3,117,708 submissions since 1991 [40]
Genomics (GWAS) "genome-wide significance threshold" of 5 × 10⁻⁸ [29] Journals
Psychology p < 0.05 — with 0.005 proposed for new discoveries across all fields [29] 36% of 100 replications significant vs 97% of originals; effects half the size [31] Journals; preregistration now common
Experimental economics p < 0.05 [29] 61% of 18; replicated effect 66% of original [32] Journals; working papers first
Social sciencein Nature and Science p < 0.05 [29] 62% of 21; effects "about 50% of the original" [33] Journals
Preclinical cancer biology p < 0.05 [29] 46% of 112; median replication effect 85% smaller [34] Journals; bioRxiv — 37,648 preprints in five years [40]
Computer science / ML A score on a shared benchmark — not a p-value at all Refereed conferences preferred: "7 months vs 1-2 years"; "4-5 evaluations per paper compared to 2-3" [39]
A dash means this deck did not verify a figure for that field — not that none exists. The rows are not comparable to each other and are not a ranking; that is the entire point. Landscape as of August 2026.

Series Artefact — Rung 4's Gate

The Cross-Field Evidence Compatibility Check


  1. What bar did this clear? "Significant" spans p < 0.05 to 3 × 10⁻⁷ — and the ASA's own third principle is that conclusions "should not be based only on whether a p-value passes a specific threshold". [29] [30]
  2. What is the unit of evidence, and who appraises it? Medicine names its standards: GRADE's four certainty levels, applied to bodies of evidence, with trials starting high and observational studies low [35] [36]; RoB 2's five domains [37]; PRISMA 2020's 27 items [38]. Ask your field-B colleague to name the equivalent — and write down the answer, including "we don't have one".
  3. What is the base rate of it holding up? 36% · 61% · 62% · 46% — four fields, four differently measured replication figures. [31] [32] [33] [34]
  4. Is the number even comparable? Citation density differs between biochemistry and mathematics by "about an order of magnitude" [41]; database coverage differs by field too [42].
  5. Did you search where field B publishes? Conferences, preprint servers and working papers are not a lesser tier everywhere. [39] [40]

The rule: if you cannot answer all five, that finding enters your memo as a question for a colleague, never as evidence — because what separates good cross-field work from poor is "how these are combined". [14]

Section 05

Team Science, and the Workflow

The cheapest cross-field tool is a colleague. Here is what AI does for a mixed team — and the four-rung ladder that holds all of today together.

Literacy Foundation

Team Science: AI Helps With Two of the Seven Problems


The seven features that create challenges for team science: "(1) high diversity of membership; (2) deep knowledge integration; (3) large size; (4) goal misalignment with other teams; (5) permeable team and group boundaries; (6) geographic dispersion; and (7) high task interdependence." National Research Council, consensus study report, 2015 [11]

Series Artefact · Today's Canonical Workflow

The Cross-Field Orientation Ladder


Rung What you hand to AI Tier · Zone The check that makes the output usable
1 · Translate→ a glossary Field B's 20 core terms — and every term it shares with field A. No papers, no citations. Tier 1 · Unreliable Verify every row against one field-B review article. Clarity is the failure mode, not the goal: AI summaries read as well as human ones and are understood less [23]
2 · Map→ a landscape One seed paper you already trust, plus your question. Not "tell me about field B". Tier 2 + a citation-graph tool · Unreliable Rebuild it in a structurally different tool — a similarity graph and a citation path fail differently, so one run is a sample, not a search. A preprint puts repeat-run overlap at 11.8–28% [28]
3 · Read→ 8 annotated papers The map, your question, and the constraint that every entry must carry a resolvable identifier. Tier 2, grounded · Unreliable Resolve all 8 in a database (Session 2's rule — they will usually all resolve now; necessary, not sufficient), then read 2 of the 8 in full yourself. If either annotation misdescribes its paper, discard the list — with valid links, factual support still runs 39–77% [66], and fabrication climbs with rarity: 6% dense, 28–29% niche [21]
4 · Synthesise→ a one-page memo Your own notes on those 8 papers — never the papers alone, and never field B's conclusions unread. Tier 1 · Dangerous as evidence Run the Compatibility Check on every borrowed finding, then have one field-B researcher read the memo. Their first correction is the value of the whole exercise
One page, one field pair, one date. You cannot skip a rung: each rung's output is the next rung's input, and each rung's check is what makes that input safe to use. Rungs are tasks, not products — the ladder survives its tools being renamed.

Series Artefact — Session 5's Canvas, Row 6

Filling In the Row You Left Empty


Lifecycle stage What you hand to AI Tier · Zone Your non-negotiable check Artefact that governs it
1–5Sessions 1–5 Frame · Discover · Read · Analyse · Write — filled in last session Tier 1–2 · mostly Unreliable As recorded on your canvas CRIT · Matrix · Triage · Guardrails · Policy Table
6 · Cross-fieldSession 6 — today Field B's vocabulary, landscape and reading list — never its conclusions Tier 1Tier 2 · Unreliable, Dangerous at rung 4 Glossary checked against one field-B review · every paper resolved · 2 of 8 read in full · one field-B reader on the memo The Cross-Field Orientation Ladder
7 · Your discipline's specialised tools — you fill this row in during Session 7
Three rules for the ladder. (1) The glossary is the deliverable, not the summary — if you cannot define field B's twenty words without the tool, you cannot audit anything downstream. (2) Two tools, or no map [27] [28]. (3) The memo is a list of questions for a colleague, not a finding — and when you publish, cite across fields but frame within one [8].

Demo Preview — One Field Pair, Twenty-Five Minutes

A Clinical Epidemiologist Gets Oriented in Machine Learning


Everything in the demo runs on eight real, DOI-verified papers, six of them open access, listed in the handout. The pairing is chosen because both fields are legible to a mixed room — and because discrimination, bias, validation and calibration mean different things on either side of it.

References

References (1–21)


  1. Uzzi, B., Mukherjee, S., Stringer, M., & Jones, B. (2013). Atypical Combinations and Scientific Impact. Science 342(6157), 468–472. doi.org/10.1126/science.1240474 — accessed 2026-07-29
  2. Larivière, V., & Gingras, Y. (2010). On the relationship between interdisciplinarity and scientific impact. JASIST 61(1), 126–131. doi.org/10.1002/asi.21226 · arxiv.org/abs/0908.1776 — accessed 2026-07-29
  3. Yegros-Yegros, A., Rafols, I., & D'Este, P. (2015). Does Interdisciplinary Research Lead to Higher Citation Impact? PLOS ONE 10(8), e0135095. doi.org/10.1371/journal.pone.0135095 — accessed 2026-07-29
  4. Wang, J., Thijs, B., & Glänzel, W. (2015). Interdisciplinarity and Impact: Variety, Balance, Disparity. PLOS ONE 10(5), e0127298. doi.org/10.1371/journal.pone.0127298 — accessed 2026-07-29
  5. Leahey, E., Beckman, C. M., & Stanko, T. L. (2017). Prominent but Less Productive. Administrative Science Quarterly 62(1), 105–139. doi.org/10.1177/0001839216665364 — accessed 2026-07-29
  6. Bromham, L., Dinnage, R., & Hua, X. (2016). Interdisciplinary research has consistently lower funding success. Nature 534, 684–687. doi.org/10.1038/nature18315 — accessed 2026-07-29
  7. Berkes, E., Marion, M., Milojević, S., & Weinberg, B. A. (2024). Slow convergence: Career impediments to interdisciplinary biomedical research. PNAS 121(32). doi.org/10.1073/pnas.2402646121 — accessed 2026-07-29
  8. Xiang, S., Romero, D. M., & Teplitskiy, M. (2025). Evaluating interdisciplinary research: Disparate outcomes for topic and knowledge base. PNAS 122(17). doi.org/10.1073/pnas.2409752122 — accessed 2026-07-29
  9. Martínez, A., & Mammola, S. (2021). Specialized terminology reduces the number of citations of scientific papers. Proc. R. Soc. B 288(1948), 20202581. doi.org/10.1098/rspb.2020.2581 — accessed 2026-07-29
  10. Vilhena, D. A., et al. (2014). Finding Cultural Holes. Sociological Science 1, 221–238. doi.org/10.15195/v1.a15 — accessed 2026-07-29
  11. National Research Council (2015). Enhancing the Effectiveness of Team Science. National Academies Press. doi.org/10.17226/19007 · ncbi.nlm.nih.gov/books/NBK310391 — accessed 2026-07-29
  12. Stirling, A. (2007). A general framework for analysing diversity in science, technology and society. J. R. Soc. Interface 4(15), 707–719. doi.org/10.1098/rsif.2007.0213 — accessed 2026-07-29
  13. Wang, Q., & Schneider, J. W. (2020). Consistency and validity of interdisciplinarity measures. Quantitative Science Studies 1(1), 239–263. doi.org/10.1162/qss_a_00011 — accessed 2026-07-29
  14. McLeish, T., & Strang, V. (2016). Evaluating interdisciplinary research: the elephant in the peer-reviewers' room. Palgrave Comms 2, 16055. doi.org/10.1057/palcomms.2016.55 — accessed 2026-07-29
  15. U.S. National Science Foundation (2026). Learn About Convergence Research. nsf.gov/funding/learn/research-types/learn-about-convergence-research — accessed 2026-07-29
  16. Hao, Q., Xu, F., Li, Y., & Evans, J. (2026). AI tools expand scientists' impact but contract science's focus. Nature 649(8099), 1237–1243. doi.org/10.1038/s41586-025-09922-y — accessed 2026-07-29
  17. Rein, D., et al. (2024). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. COLM 2024. arxiv.org/abs/2311.12022 — accessed 2026-07-29
  18. Wang, Y., Ma, X., Zhang, G., et al. (2024). MMLU-Pro. NeurIPS 2024 Datasets & Benchmarks. arxiv.org/abs/2406.01574 — accessed 2026-07-29
  19. Phan, L., et al. (2026). A benchmark of expert-level academic questions to assess AI capabilities (HLE). Nature 649(8099), 1139–1146. doi.org/10.1038/s41586-025-09962-4 — accessed 2026-07-29
  20. Kang, Z., et al. (2026). HSSBench: Humanities and Social Sciences Ability for Multimodal LLMs. Preprint. arxiv.org/abs/2506.03922 — accessed 2026-07-29
  21. Linardon, J., et al. (2025). Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication. JMIR Mental Health 12, e80371. doi.org/10.2196/80371 — accessed 2026-07-29

References

References (22–41)


  1. Peters, U., & Chin-Yee, B. (2025). Generalization bias in large language model summarization of scientific research. R. Soc. Open Sci. 12(4), 241776. doi.org/10.1098/rsos.241776 — accessed 2026-07-29
  2. Guo, Y., Sohn, J. H., Leroy, G., & Cohen, T. (2026). Are LLM-generated plain language summaries truly understandable? J. Biomed. Inform. 179, 105038. doi.org/10.1016/j.jbi.2026.105038 — accessed 2026-07-29
  3. Goldsack, T., Scarton, C., Shardlow, M., & Lin, C. (2024). Overview of the BioLaySumm 2024 Shared Task. BioNLP @ ACL 2024. aclanthology.org/2024.bionlp-1.10/ — accessed 2026-07-29
  4. Lewis, M., & Mitchell, M. (2024). Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in LLMs. Preprint. arxiv.org/abs/2402.08955 · arxiv.org/abs/2411.14215 — accessed 2026-07-29
  5. Shen, Y., et al. (2026). IDRBench: Understanding the Capability of LLMs on Interdisciplinary Research. Preprint. arxiv.org/abs/2507.15736 — accessed 2026-07-29
  6. Wright, D., et al. (2026). Epistemic Diversity and Knowledge Collapse in Large Language Models. Preprint. arxiv.org/abs/2510.04226 — accessed 2026-07-29
  7. Dathe, A., Hoffmann, K., & Mangold, A. (2026). Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research. Preprint. arxiv.org/abs/2605.10125 — accessed 2026-07-29
  8. Benjamin, D. J., et al. (2018). Redefine statistical significance. Nature Human Behaviour 2(1), 6–10. doi.org/10.1038/s41562-017-0189-z — accessed 2026-07-29
  9. Wasserstein, R. L., & Lazar, N. A. (2016). The ASA Statement on p-Values. The American Statistician 70(2), 129–133. doi.org/10.1080/00031305.2016.1154108 · amstat.org — accessed 2026-07-29
  10. Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science 349(6251), aac4716. doi.org/10.1126/science.aac4716 — accessed 2026-07-29
  11. Camerer, C. F., et al. (2016). Evaluating replicability of laboratory experiments in economics. Science 351(6280), 1433–1436. doi.org/10.1126/science.aaf0918 — accessed 2026-07-29
  12. Camerer, C. F., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science. Nature Human Behaviour 2, 637–644. doi.org/10.1038/s41562-018-0399-z — accessed 2026-07-29
  13. Errington, T. M., et al. (2021). Investigating the replicability of preclinical cancer biology. eLife 10, e71601. doi.org/10.7554/eLife.71601 — accessed 2026-07-29
  14. GRADE Working Group — Schünemann, H., et al. (Eds.) (2013). GRADE Handbook. gdt.gradepro.org/app/handbook/handbook.html — accessed 2026-07-29
  15. Balshem, H., et al. (2011). GRADE guidelines: 3. Rating the quality of evidence. J. Clin. Epidemiol. 64(4), 401–406. doi.org/10.1016/j.jclinepi.2010.07.015 — accessed 2026-07-29
  16. Sterne, J. A. C., et al. (2019). RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 366, l4898. doi.org/10.1136/bmj.l4898 — accessed 2026-07-29
  17. Page, M. J., et al. (2021). The PRISMA 2020 statement. BMJ 372, n71. doi.org/10.1136/bmj.n71 — accessed 2026-07-29
  18. Patterson, D., Snyder, L., & Ullman, J. (1999). Evaluating Computer Scientists and Engineers For Promotion and Tenure. CRA Best Practices Memo. cra.org/resources/best-practice-memos — accessed 2026-07-29
  19. arXiv (2026). Monthly submission statistics; and Abdill, R. J., & Blekhman, R. (2019). Tracking the popularity and outcomes of all bioRxiv preprints. eLife 8, e45133. arxiv.org/stats/monthly_submissions · doi.org/10.7554/eLife.45133 — accessed 2026-07-29
  20. Waltman, L. (2016). A review of the literature on citation impact indicators. J. Informetrics 10(2), 365–391. doi.org/10.1016/j.joi.2016.02.007 · arxiv.org/abs/1507.02099 — accessed 2026-07-29

References

References (42–66)


  1. Martín-Martín, A., et al. (2021). …a multidisciplinary comparison of coverage via citations. Scientometrics 126(1), 871–906. doi.org/10.1007/s11192-020-03690-4 — accessed 2026-07-29
  2. Connected Papers (2026). About; Pricing. connectedpapers.com/about · /pricing — accessed 2026-07-29
  3. Inciteful (2026). Paper Discovery; Literature Connector; Data Sources. incitefulmed.com/academic/ (formerly inciteful.xyz) — accessed 2026-07-29
  4. ResearchRabbit (2026). Home; Pricing; Help guide. researchrabbit.ai · /pricing · /help/guide — accessed 2026-07-29
  5. Litmaps (2026). Pricing; Documentation. litmaps.com/pricing · docs.litmaps.com — accessed 2026-07-29
  6. Open Knowledge Maps (2026). About; FAQ. openknowledgemaps.org/about · /faq — accessed 2026-07-29
  7. VOSviewer (2026). Download; Features. v1.6.21, 12 June 2026. vosviewer.com/download · /features/highlights — accessed 2026-07-29
  8. OpenAlex (2026). Topics; Semantic search; Authors; Authentication & pricing — plus live API counts. developers.openalex.org · api.openalex.org — accessed 2026-07-29
  9. Semantic Scholar (2026). Academic Graph API; Recommendations API. api.semanticscholar.org/api-docs/graph · /recommendations — accessed 2026-07-29
  10. Elsevier (2024/2026). What is the complete list of ASJC subject areas in Scopus? Scopus Support Center. service.elsevier.com/app/answers/detail/a_id/15181 — accessed 2026-07-29
  11. Clarivate (2025). Web of Science Core Collection: Web of Science Categories. support.clarivate.com — accessed 2026-07-29
  12. Zotero (2026). Groups; Storage. zotero.org/support/groups · zotero.org/storage — accessed 2026-07-29
  13. Google (2026). Gemini Notebook (formerly NotebookLM) — FAQ; Upgrade your plan. support.google.com/gemininotebook/answer/16269187 · /16213268 — accessed 2026-07-29
  14. NIH (2026). NIH RePORTER APIs. api.reporter.nih.gov — accessed 2026-07-29
  15. ORCID (2026). What is ORCID? info.orcid.org/what-is-orcid/ — accessed 2026-07-29
  16. Collins, G. S., et al. (2024). TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385, e078378. doi.org/10.1136/bmj-2023-078378 — accessed 2026-07-29
  17. Van Calster, B., et al. (2019). Calibration: the Achilles heel of predictive analytics. BMC Medicine 17, 230. doi.org/10.1186/s12916-019-1466-7 — accessed 2026-07-29
  18. Obermeyer, Z., et al. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464), 447–453. doi.org/10.1126/science.aax2342 — accessed 2026-07-29
  19. Christodoulou, E., et al. (2019). A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. J. Clin. Epidemiol. 110, 12–22. doi.org/10.1016/j.jclinepi.2019.02.004 — accessed 2026-07-29
  20. Assel, M., & Vickers, A. (2025). The F score ranks diagnostic tests and prediction models inconsistently with their clinical utility. Diagn. Progn. Res. 9(1), 30. doi.org/10.1186/s41512-025-00214-7 — accessed 2026-07-29
  21. Paulus, J. K., & Kent, D. M. (2020). Predictably unequal: … algorithmic clinical prediction may increase health disparities. npj Digital Medicine 3, 99. doi.org/10.1038/s41746-020-0304-9 — accessed 2026-07-29
  22. Sung, J., & Hopper, J. L. (2023). Co-evolution of epidemiology and artificial intelligence: challenges and opportunities. Int. J. Epidemiol. 52(4), 969–973. doi.org/10.1093/ije/dyad089 — accessed 2026-07-29
  23. Peters, U., et al. (2026). Generics in science communication. Public Underst. Sci., online 20 Apr 2026. doi.org/10.1177/09636625261425891 — accessed 2026-08-21
  24. Artificial Analysis (2026). MMLU-Pro and Humanity's Last Exam evaluation pages (independent benchmark runs). artificialanalysis.ai/evaluations/mmlu-pro · /humanitys-last-exam — accessed 2026-08-21
  25. Onweller, H., et al. (2026). Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents. Preprint. arxiv.org/abs/2605.06635 — accessed 2026-08-21