← series index
# Session 6: AI for Multidisciplinary Research
*AI for Researchers — attendee handout · Landscape as of August 2026*
---
## At a glance
| | |
|---|---|
| **Session** | 6 of 7 — AI for Multidisciplinary Research |
| **You'll learn** | Why crossing a field boundary is measurably hard, what AI is genuinely good at across that boundary, where it fails hardest — and a four-rung workflow that gets you oriented in an unfamiliar field without being fooled. |
| **You'll practise** | All four rungs live: a glossary with false friends flagged, a landscape map built in two structurally different tools, an eight-paper annotated reading list with every DOI resolved, and a cross-field synthesis memo. |
| **Prerequisite** | None. Uses Session 1's CRIT and zones and Session 2's discovery rules; fills **row 6** of Session 5's canvas. |
| **Tools referenced** | None endorsed. The Ladder names tasks and checks, so nothing here becomes wrong when a product is renamed. |
---
## The one-paragraph version
The prize is real: across 17.9 million papers, the highest-impact work pairs a conventional core with an intrusion of unusual combinations, and such papers were **twice as likely to be highly cited** [1]. So is the penalty. Across **all 18,476** proposals to one national scheme, "the greater the degree of interdisciplinarity, the lower the probability of being funded" [6]. Across **154,021** biomedical PhDs, it takes about **8 years** for half of the most interdisciplinary cohort to stop publishing, against **more than 20 years** for moderately interdisciplinary researchers [7]. The mechanism is language, and then epistemology: jargon in a title and abstract measurably reduces citations [9], and disciplines' "unique languages reflect deeper differences in underlying assumptions, epistemologies" [11]. AI helps here — but it is weakest exactly where you are weakest. In one 2025 experimental study, AI-generated citations were fabricated in **6%** of cases for a densely published topic and **28–29%** for two niche ones — same model, same task [21]. Grounded 2026 assistants change the absolute rates, not the logic: in a May-2026 audit of 14 current models, the strongest kept link validity above 94% while factual support for the cited claims ran only **39–77%** [66] — the risk has moved from "does the paper exist" to "does it say that". Its most seductive service is the one that misleads: AI plain-language summaries were rated as clear as human-written ones while producing **significantly worse comprehension** [23], and across 4,900 summaries from the 2024–25 model cohort, models overgeneralised in **26–73%** of cases — with prompting for accuracy roughly *doubling* the problem [22]; the same team's 2026 follow-up finds GPT-5-class models still rate generic claims as more generalizable than laypeople do, where domain experts rate them as less [64]. The Ladder below is built around that: four rungs, four outputs, four checks, and you cannot skip a rung.
---
## The Cross-Field Orientation Ladder — checklist and prompt templates
Adapt the bracketed parts. `[FIELD A]` is yours; `[FIELD B]` is theirs.
**Rung 1 · Translate → a glossary.** Zone: Unreliable. Tier 1.
```
Context: I am a [FIELD A] researcher with no training in [FIELD B]. I need
to read [FIELD B] papers on [TOPIC] well enough to judge whether their
evidence bears on my question.
Role and register: Act as a methodologist who works in both fields.
Assume strong general research literacy; assume no [FIELD B] knowledge.
Instructions and constraints:
- Give the 20 terms I cannot read a [FIELD B] abstract without; one
plain sentence each.
- Then FLAG EVERY TERM THAT ALSO EXISTS IN [FIELD A] WITH A DIFFERENT
MEANING, both meanings side by side.
- Then list terms where the two fields use DIFFERENT WORDS FOR THE SAME
THING.
- Mark anything you are unsure of "LOW CONFIDENCE".
- Do not cite any paper and do not name any author.
- Do not use web search for this step; I will ground the result myself.
Task: Produce the glossary, the false-friend table, the synonym table.
```
- [ ] **Verify every row against one field-B review article or reporting guideline.** Ten minutes.
- [ ] *Watch out for:* fluency. Clarity is the failure mode here, not the goal [22][23].
**Rung 2 · Map → a landscape.** Zone: Unreliable. Tier 2 plus a citation-graph tool.
```
Context: I am mapping [FIELD B] as an outsider from [FIELD A]. My seed
paper is [CITATION + DOI], which I have read.
Instructions and constraints:
- Name the 4-6 sub-communities of [FIELD B] working on [TOPIC], and what
distinguishes each one's method.
- Name the venues each publishes in, and say whether they are journals,
conferences or preprint servers.
- Name the 5 papers everyone in this area cites, with DOIs. If you are
not certain of a DOI, write "DOI UNVERIFIED".
- Flag any sub-community whose evidence standards differ from [FIELD A]'s.
Task: Produce the map, then list what you are least confident about.
```
- [ ] **Rebuild the map in a structurally different tool** — a similarity graph and a shortest-citation-path fail in different directions, so one run is a sample, not a search. A *preprint* puts repeat-run overlap at **11.8–28%** [28].
- [ ] *Watch out for:* asking a chatbot alone. A *preprint* testing 27 models (through the 2025 generation) found "all models are less epistemically diverse than a basic web search" [27]; the August-2026 cohort is untested.
**Rung 3 · Read → 8 annotated papers.** Zone: Unreliable. Tier 2, grounded.
```
Instructions and constraints:
- Propose exactly 8 papers to orient me, ranked in reading order.
- For each: full citation, DOI, two sentences on why it is on the list,
one sentence on what I can safely skip, and open-access status.
- Prefer reporting guidelines, systematic reviews and methodological
critiques over individual results papers.
- If you are not certain a DOI is correct, write "DOI UNVERIFIED".
Task: Produce the ranked, annotated reading list as a table.
```
- [ ] **Resolve all eight in a database before use** (Session 2's rule). Expect them all to resolve on a current live-search assistant — resolving is necessary, not sufficient: factual support runs **39–77%** even where links are valid [66].
- [ ] **Read two of the eight in full** and check them against their annotations. This is now the real check. If either is misdescribed, **discard the list** — you no longer know which of the other six are wrong.
**Rung 4 · Synthesise → a one-page memo.** Zone: **Dangerous** as evidence. Tier 1, on *your own notes*.
```
Context: Below are MY OWN notes on 8 papers I have read about [FIELD B].
[PASTE NOTES]
Instructions and constraints:
- One page, for a colleague in [FIELD A], not in [FIELD B].
- Structure: (1) what field B claims; (2) what its own critics say;
(3) the three terms where our fields mean different things; (4) what
I would need to see before believing a result from field B.
- Every claim must be attributable to one of MY notes. If it is not in
my notes, do not make it.
- Where the fields hold different standards of evidence, SAY SO rather
than reconciling them.
- End with three questions to ask a collaborator from field B.
Task: Write the memo.
```
- [ ] **Run the Compatibility Check (below) on every borrowed finding.**
- [ ] **Have one field-B researcher read it.** Their first correction is the value of the whole exercise.
---
## The Cross-Field Evidence Compatibility Check
Before any field-B finding enters your evidence base, answer all five. **If you cannot, it goes in the memo as a question, never as a finding.**
1. **What bar did it clear?** "Significant" spans *p* < 0.05, the genome-wide 5 × 10⁻⁸, and particle physics' 5-sigma ≈ 3 × 10⁻⁷ — all named in one paper [29]. The ASA's own third principle: conclusions "should not be based only on whether a p-value passes a specific threshold" [30].
2. **What is the unit of evidence, and who appraises it?** Medicine names its standards: GRADE's four certainty levels, applied to *bodies* of evidence with trials starting high and observational studies low [35][36]; RoB 2's five domains [37]; PRISMA 2020's 27 items [38]. Ask your field-B colleague to name the equivalent — and write down the answer, including "we don't have one".
3. **What is the base rate of it holding up?** Four fields, four differently measured replication figures: **36%** of 100 in psychology [31]; **61%** of 18 in experimental economics [32]; **62%** of 21 in social science [33]; **46%** of 112 effects in preclinical cancer biology, with the median replication effect **85% smaller** [34].
4. **Is the number even comparable?** Citation density differs between biochemistry and mathematics by "about an order of magnitude" [41]; database coverage differs by field, with gaps clustering "especially in the Humanities" [42].
5. **Did you search where field B publishes?** Computer science prefers refereed conferences — "7 months vs 1-2 years" to print, and "4-5 evaluations per paper compared to 2-3" [39]. arXiv has taken **3,117,708** submissions since 1991; bioRxiv took **37,648** in its first five years [40].
---
## Cross-field discovery tools
Every cell from that tool's own documentation, fetched 2026-07-29. Rows are **capabilities**, not endorsements.
| Tool | Best for | Access / cost | Key limitation |
|---|---|---|---|
| **Connected Papers** | A similarity graph around one seed you trust | Free: 5 graphs/month; Academic USD3/mo | "**Connected Papers is not a citation tree**" — adjacency means similarity, from co-citation and bibliographic coupling [43] |
| **Inciteful** (Literature Connector) | The cross-field move: shortest citation paths **between two literatures** | "100% Free to use"; 240M+ papers [44] | Shortest paths only, one paper per end, max 6 hops. Note: `inciteful.xyz` now redirects to `incitefulmed.com/academic/` [44] |
| **ResearchRabbit** | Similar / Earlier / Later work from seeds; 310M+ papers | "$0, Forever!" — capped at **50 seed articles**; RR+ $10/mo → 300 [45] | Upstream corpus provider not documented [45] |
| **Litmaps** | Multi-seed maps of papers that "either cite or are cited by" your seeds | Free: 2 maps, 100 articles each; Pro USD10/mo [46] | Ranking behind "Discover" not documented [46] |
| **Open Knowledge Maps** | Mapping a *query* rather than a seed paper | Free, non-profit, open source; PubMed or BASE [47] | **100 documents per map**, by design [47] |
| **OpenAlex semantic search** | Paste a paragraph, not a phrase; retrieves by meaning | Usage-priced since 2026: "$1 of free usage every day" with a key [49] | 2,000-char input, 50 results, 1 req/s — and browsing the website spends the same budget [49] |
| **Semantic Scholar Recommendations** | "Like these, unlike those" — takes **negative** examples too | Free; most endpoints need no key [50] | Recommendation pool is only "recent" or "all-cs" [50] |
| **OpenAlex Authors · ORCID · NIH RePORTER** | Finding *people* in field B — funded projects are earlier signal than papers | ORCID free; RePORTER "available to all public users" [49][55][56] | RePORTER is US federal awards only, one request per second [55] |
*Landscape as of August 2026; capabilities and pricing change monthly — verify before relying on a cell.*
---
## Pitfalls
**False friends — the same word, changed meaning.** Worked here on one real pair, clinical prediction ↔ machine learning; build your own before rung 2.
| Word | In clinical prediction | In machine learning |
|---|---|---|
| **Discrimination** | Desirable: "how well the predictions from the model differentiate between individuals with and without the outcome" [57][58] | Undesirable: "unfair prejudice leading to inequity" [62] — a model can have excellent discrimination *and* be discriminatory |
| **Bias** | "Systematic errors… which arise from weaknesses in study designs and conduct" [63] | "In AI, bias means the intercepts in unit models" [63] — plus the fairness sense [59]. Three meanings, one word |
| **Validation** | External validation. So contested that TRIPOD+AI states "we refer to **validation as evaluation** in this article" [57] | "data used for parameter tuning" [57]; field A's *external* validation is field B's **test** set [63] |
| **Calibration** | "Agreement between observed outcomes and estimated values" — routinely required [57][58] | Often simply absent: not addressed in **56 of 71 (79%)** ML-vs-regression studies [60] — a *missing* concept, the hardest gap to notice |
| **Sensitivity · PPV** | Sensitivity; positive predictive value | Recall; precision — "what are more typically described in medical contexts as positive predictive value and sensitivity" [61] |
**The other four pitfalls.**
- **Fluency mistaken for understanding.** Rate your glossary by whether you can source it, never by how well it reads [22][23].
- **A finding travelling without its conditions.** Every borrowed number carries a design, a threshold and a population. Move all three or move none — that is what the Compatibility Check is for.
- **One tool, one run.** A *preprint* measured precision at **21–41%** and run-to-run overlap at **11.8–28%** across five tools [28].
- **Framing the paper the way you researched it.** On 128,950 manuscripts including rejections, interdisciplinary *references* raised acceptance while interdisciplinary *topic language* lowered it [8]. **Cite across fields; frame within one.**
---
## Row 6 of your Personal AI Workflow Canvas
| Lifecycle stage | What I hand to AI | Tier · Zone | My non-negotiable check | Artefact |
|---|---|---|---|---|
| **6 · Cross-field** (S6) | Field B's vocabulary, landscape and reading list — *never* its conclusions | Tier ___ · ☐ Safe ☐ Unreliable ☐ Dangerous | ☐ glossary checked against one field-B review ☐ every paper resolved ☐ 2 of 8 read in full ☐ one field-B reader on the memo | **The Cross-Field Orientation Ladder** |
**Field A:** ______________ **Field B:** ______________ **Date filled:** __________ **My field-B reader:** ______________
---
## Further reading
- Collins, G. S., et al. (2024). *TRIPOD+AI statement*. BMJ 385:e078378. https://doi.org/10.1136/bmj-2023-078378 — read Box 1 even if you never build a prediction model. It is what a rung-1 glossary looks like when a field writes one down.
- Sung, J., & Hopper, J. L. (2023). *Co-evolution of epidemiology and artificial intelligence*. Int J Epidemiol 52(4):969–973. https://doi.org/10.1093/ije/dyad089 — five pages, one translation table, three meanings of "bias". The model for your own pair.
- Berkes, E., et al. (2024). *Slow convergence: Career impediments to interdisciplinary biomedical research*. PNAS 121(32). https://doi.org/10.1073/pnas.2402646121 — the 8-versus-20-years finding; worth showing a head of department.
- Hao, Q., Xu, F., Li, Y., & Evans, J. (2026). *AI tools expand scientists' impact but contract science's focus*. Nature 649:1237–1243. https://doi.org/10.1038/s41586-025-09922-y — the tension this session sits inside.
- Bzdok, D., Altman, N., & Krzywinski, M. (2018). *Statistics versus machine learning*. Nature Methods 15:233–234. https://doi.org/10.1038/nmeth.4642 — two pages on inference versus prediction; the gentlest on-ramp to a neighbouring field.
- National Research Council (2015). *Enhancing the Effectiveness of Team Science*. https://doi.org/10.17226/19007 — Chapter 2's seven features diagnose why your last cross-field project was hard.
*Full annotated source list — 66 sources, with the verified wording and the caveat behind every number: see `sources.md`.*
---
*AI for Researchers · Session 6: AI for Multidisciplinary Research · Landscape as of August 2026*