AI for Researchers
Getting oriented in a field that is not yours — in four rungs, with the checks that keep it honest
{{Presenter Name}}
Landscape as of August 2026
2-Minute Recap
Today (Session 6): the job no colleague has time to do for you — getting oriented in someone else's field, fast enough to be useful and checked enough to be safe.
Learning Objectives
Objectives 4–6 continue on the next slide.
Learning Objectives
Section 01
Not a confidence problem, and not a you problem. A measured, structural one — and knowing its exact shape tells you what to delegate.
Literacy Foundation — Start With the Prize
The Evidence — Measured, Not Asserted
Read with the caveat its own field states: interdisciplinarity indicators give "surprisingly deviant results" and "should be interpreted with much caution". [13] The direction is consistent across studies; the effect sizes are not. Practical reading of [8]: cite across fields, frame within one.
Critical Literacy — The Key Slide
That last number is about humans, not models. Thirty minutes of searching does not make you a non-expert who can read field B — it makes you a non-expert with tabs open. This session is about the thirty minutes that does work.
Section 02
Rung 1 of today's workflow — the one thing AI is unusually good at, and the one place it will mislead you most convincingly.
Literacy Foundation
A translation you cannot audit is not a translation. It is a feeling of comprehension — and the next two slides are about how reliably AI produces that feeling without the comprehension.
The Terminology Trap — Worked on One Real Field Pair
| The word | In clinical prediction research | In machine learning | Why it bites |
|---|---|---|---|
| Discrimination | Desirable. "How well the predictions… differentiate between individuals with and without the outcome" [57] [58] | Undesirable. "Unfair prejudice leading to inequity" [62] [59] | A model can have excellent discrimination and be discriminatory — both true at once |
| Bias | "Systematic errors… which arise from weaknesses in study designs and conduct" [63] | "In AI, bias means the intercepts in unit models" [63] — plus the fairness sense [59] | Three meanings, one word. "Low bias" is praise in one, and silent on the other two |
| Validation | External validation — so contested that TRIPOD+AI "refer[s] to validation as evaluation" throughout [57] | "data used for parameter tuning" [57]; epi's external validation is ML's test set [63] | "We validated the model" names two very different amounts of evidence |
| Calibration | "Agreement between observed outcomes and estimated values" — routinely required [57] [58] | Often absent: not addressed in 56 of 71 (79%) ML-vs-regression studies [60] | Not a different meaning — a missing concept. The hardest kind of gap to notice |
| Sensitivity= recall | Sensitivity: of those with the outcome, the share the model flags | Recall — one of the terms "more typically described in medical contexts as positive predictive value and sensitivity" [61] | Same quantity, two names. Search one, find nothing |
| PPV= precision | Positive predictive value: of those flagged, the share that have it | Precision — "more commonly termed positive predictive value" [61] | The pair that derails a methods conversation fastest |
The Evidence — Two Results That Should Change How You Prompt
So the check on rung 1 is not "does this read clearly?" — clarity is the failure mode. It is: can I find each of these twenty definitions in one field-B review article?
Critical Literacy — Session 1's Rule 2, Quantified
The live per-discipline evidence is now Humanity's Last Exam: "low accuracy and calibration" at the expert frontier [19]; best tracked score 55.5% (Claude Fable 5, Artificial Analysis, Aug 2026) [65] — better, still far from expert level. [20], [25] and [26] are preprints on pre-2026 cohorts. It is all Session 1's Rule 2: rarity moves tasks rightward.
CRIT, Applied — Rung 1
"No papers, no web search" is deliberate. Rung 1 is a vocabulary task; with retrieval on by default, the model will otherwise reach for pages you cannot audit yet — and citations are Session 1's Dangerous zone: [21] is why. Papers arrive on rung 2, from a database, or not at all.
Section 03
Rung 2 — where AI stops being the source and becomes the index. Three search geometries, and one question about who decided what field your paper is in.
Literacy Foundation
So an interdisciplinary paper in a disciplinary journal inherits its venue's fields — invisible to a field-B browse. Coverage differs too: gaps cluster "especially in the Humanities". [42]
Series Artefact — Extending Session 2's Matrix
| Tool | What it maps, and from what | Access, at last check | A limit it documents itself |
|---|---|---|---|
| Connected Papersone seed paper | A similarity graph from co-citation + bibliographic coupling; "~50,000 papers" analysed per graph [43] | Free: 5 graphs/month. Academic USD3/mo [43] | "Connected Papers is not a citation tree" — adjacency is similarity, not citation [43] |
| IncitefulLiterature Connector | Two seeds: "how two bodies of literature connect… through citations". 240M+ papers [44] | "100% Free to use"; on OpenAlex, S2, Crossref, OpenCitations [44] | Shortest paths only; one paper per end; max 6 hops [44] |
| ResearchRabbitcarried from Session 2 | Similar / Earlier / Later work and author networks; 310M+ papers [45] | "$0, Forever!", capped at 50 seeds. RR+ $10/mo → 300 [45] | Upstream corpus provider not documented [45] |
| Litmapsmulti-seed | Papers that "either cite or are cited by" the seed; 270M+ articles [46] | Free: 2 maps, 100 articles. Pro USD10/mo [46] | Ranking behind "Discover" not documented [46] |
| Open Knowledge Mapsno seed paper needed | A query, clustered from a word co-occurrence matrix over titles, abstracts, keywords [47] | Free; non-profit; open source. PubMed or BASE [47] | 100 documents per map, by design — "a manageable amount" [47] |
| VOSviewerbring your own export | Co-citation, coupling and term co-occurrence networks from a WoS / Scopus / PubMed export [48] | "can be used freely for any purpose"; v1.6.21, June 2026 [48] | "Free to use" — the site names no open-source licence, so nor does this deck [48] |
Series Artefact — Extending Session 2's Matrix
| Capability | What it does, documented | Access, at last check | A limit it documents itself |
|---|---|---|---|
| OpenAlex semantic searchpaste a paragraph, not a phrase | Embeds every title + abstract into a 1,024-dim vector: "the richer the input, the better the matches" [49] | Usage-priced since 2026: "$1 of free usage every day" with a key [49] | 2,000-char input, 50 results, 1 req/s — and browsing the website spends the same budget [49] |
| Semantic Scholar Recommendationsthe "unlike this" lever | Takes positive and negative example papers — the documented way to push results away from your field [50] | Free; most endpoints need no key. 214M papers [50] | Pool is only "recent" or "all-cs"; the spec does not say it uses SPECTER2 [50] |
| OpenAlex topic filtersmove one level up | Every work carries a primary_topic with its full path: 4 domains / 26 fields / ~252 subfields / ~4,500 topics [49] |
Free key tier as above; List+Filter is the cheapest class [49] | The classification model is not documented — the docs say only "automatically assigned" [49] |
| OpenAlex Authors · ORCIDwho works in field B | Filter author profiles by topic, institution, country or ORCID [49] [56] | ORCID: "free of charge to researchers" [56] | Author queries are billable List+Filter calls [49] [56] |
| NIH RePORTERwho is funded right now | Awards "from both NIH and non-NIH federal agencies" — earlier signal than papers [55] | "available to all public users"; free; bulk download [55] | US federal awards only; "no more than one URL request per second" [55] |
Critical Literacy
[27] and [28] are preprints; [16] is peer-reviewed in Nature. Read together they are the argument for this whole session: AI makes you faster at crossing fields, and makes science narrower — unless you use it to reach the people and papers it would otherwise hide.
Section 04
Rungs 3 and 4 — where "significant", "validated" and "replicated" turn out to name seven different things, and combining them naively is the failure mode.
The Evidence
| Field | The threshold convention | A published replication figure | Where the work appears |
|---|---|---|---|
| Particle physics | 5-sigma — "roughly a P value threshold of 3 × 10⁻⁷" [29] | — | Journals and arXiv: 3,117,708 submissions since 1991 [40] |
| Genomics (GWAS) | "genome-wide significance threshold" of 5 × 10⁻⁸ [29] | — | Journals |
| Psychology | p < 0.05 — with 0.005 proposed for new discoveries across all fields [29] | 36% of 100 replications significant vs 97% of originals; effects half the size [31] | Journals; preregistration now common |
| Experimental economics | p < 0.05 [29] | 61% of 18; replicated effect 66% of original [32] | Journals; working papers first |
| Social sciencein Nature and Science | p < 0.05 [29] | 62% of 21; effects "about 50% of the original" [33] | Journals |
| Preclinical cancer biology | p < 0.05 [29] | 46% of 112; median replication effect 85% smaller [34] | Journals; bioRxiv — 37,648 preprints in five years [40] |
| Computer science / ML | A score on a shared benchmark — not a p-value at all | — | Refereed conferences preferred: "7 months vs 1-2 years"; "4-5 evaluations per paper compared to 2-3" [39] |
Series Artefact — Rung 4's Gate
The rule: if you cannot answer all five, that finding enters your memo as a question for a colleague, never as evidence — because what separates good cross-field work from poor is "how these are combined". [14]
Section 05
The cheapest cross-field tool is a colleague. Here is what AI does for a mixed team — and the four-rung ladder that holds all of today together.
Literacy Foundation
Series Artefact · Today's Canonical Workflow
| Rung | What you hand to AI | Tier · Zone | The check that makes the output usable |
|---|---|---|---|
| 1 · Translate→ a glossary | Field B's 20 core terms — and every term it shares with field A. No papers, no citations. | Tier 1 · Unreliable | Verify every row against one field-B review article. Clarity is the failure mode, not the goal: AI summaries read as well as human ones and are understood less [23] |
| 2 · Map→ a landscape | One seed paper you already trust, plus your question. Not "tell me about field B". | Tier 2 + a citation-graph tool · Unreliable | Rebuild it in a structurally different tool — a similarity graph and a citation path fail differently, so one run is a sample, not a search. A preprint puts repeat-run overlap at 11.8–28% [28] |
| 3 · Read→ 8 annotated papers | The map, your question, and the constraint that every entry must carry a resolvable identifier. | Tier 2, grounded · Unreliable | Resolve all 8 in a database (Session 2's rule — they will usually all resolve now; necessary, not sufficient), then read 2 of the 8 in full yourself. If either annotation misdescribes its paper, discard the list — with valid links, factual support still runs 39–77% [66], and fabrication climbs with rarity: 6% dense, 28–29% niche [21] |
| 4 · Synthesise→ a one-page memo | Your own notes on those 8 papers — never the papers alone, and never field B's conclusions unread. | Tier 1 · Dangerous as evidence | Run the Compatibility Check on every borrowed finding, then have one field-B researcher read the memo. Their first correction is the value of the whole exercise |
Series Artefact — Session 5's Canvas, Row 6
| Lifecycle stage | What you hand to AI | Tier · Zone | Your non-negotiable check | Artefact that governs it |
|---|---|---|---|---|
| 1–5Sessions 1–5 | Frame · Discover · Read · Analyse · Write — filled in last session | Tier 1–2 · mostly Unreliable | As recorded on your canvas | CRIT · Matrix · Triage · Guardrails · Policy Table |
| 6 · Cross-fieldSession 6 — today | Field B's vocabulary, landscape and reading list — never its conclusions | Tier 1 → Tier 2 · Unreliable, Dangerous at rung 4 | Glossary checked against one field-B review · every paper resolved · 2 of 8 read in full · one field-B reader on the memo | The Cross-Field Orientation Ladder |
| 7 · Your discipline's specialised tools — you fill this row in during Session 7 | ||||
Demo Preview — One Field Pair, Twenty-Five Minutes
Everything in the demo runs on eight real, DOI-verified papers, six of them open access, listed in the handout. The pairing is chosen because both fields are legible to a mixed room — and because discrimination, bias, validation and calibration mean different things on either side of it.
References
References
References