← series index
# AI for Researchers — Series Status
*Delivery-facing summary · compiled 2026-08-08, revised 2026-09-02 after the model-generation modernization pass · Landscape as of August 2026*
Seven sessions, four artefacts each, plus a master curriculum and shared templates. Everything is written and sourced. What remains is the post-modernization render re-verification (§1a) and presenter preparation.
---
## 0. The 2026-09 modernization pass (what changed since 2026-08-08)
The whole series was modernized from the GPT-4o / o1 / Claude 3.5–3.7-era evidence base to the August-2026 model generation (Claude Fable 5/5.1, Claude Opus 5, GPT-5.6 Sol/Terra/Luna, Kimi K3, Gemini 3.1 Pro / 3.7 Flash, DeepSeek V4). Governing documents: `modernization-notes/rewrite-spec.md` (rules), the four `modernization-notes/inventory-*.md` files (line-by-line change lists), and the four `modernization-notes/research-*.md` reports (verified evidence with URLs and accessed dates). Highlights:
- **Old-model findings were historicized, not erased**: era labels sit next to every 2022–25 number (GPT-3.5/4 fabrication rates, Liang detector bias, METR slowdown, Gartlehner extraction, silicon-samples debate), paired with current-generation companions where verified evidence exists.
- **New current-generation evidence added throughout**: Opus 5 / Fable 5–Mythos 5 system-card hallucination tradeoffs (S1/S3), the Onweller 14-model deep-research audit — link validity >94% vs 39–77% factual support (S2/S6), the 2026 benchmark trend pairs — CORE-Bench 21%→77.8%, MLE-bench 16.9%→64.44%, SciCode →60.2%, AstaBench 58%/~3% end-to-end, StatABench 68.6% (S5), HLE as MMLU-Pro's successor (S6), the Peters & Chin-Yee 2026 follow-up (S3/S6), the August-2026 detector-evasion study and the arXiv one-year ban (S4), Siler PNAS 2026 paired with He & Bu (S4).
- **He & Bu figures verified correct and precisely scoped** ("three figures, three windows, one study": ~0.1% cumulative · 0.43% 2025-Q1 quarterly rate · 40:1 use-to-disclosure ratio) in S4 and `curriculum.md`.
- **Every demo script converted from "expect the model to fail" to dual-path checks** — current models often no longer fabricate/mispool/fail-to-compile on cue; the failure branches remain as alternates, the checks are the lesson (S1 citation trap, S5 Simpson's-paradox beat, S6 DOI misses, S7 Lean kernel demo inverted to the false-theorem primary path).
- **Verification pedagogy updated**: "does the link resolve" demoted from sufficient to necessary everywhere; "does the source support the claim" is the load-bearing check.
- Landscape stamps → **August 2026**; accessed dates were updated *only* on sources actually re-fetched (2026-08-21 by research agents, 2026-09-02 by rewriters/coordinator); untouched sources honestly keep their July dates.
- `curriculum.md` re-synced: Session 1 objectives 1–2 rewritten (extended thinking is default-on), Session 5 abstract/outline moved to present-tense agentic framing, Session 4 disclosure figures scope-tightened. All 41 objectives verified verbatim-identical between curriculum and decks on 2026-09-02.
---
## 1. What exists
Project root: `/home/ehsan/reports/ai-research/`
| Session | Directory | slides | sources | handout (words) | demo script (words) |
|---|---|---|---|---|---|
| 1 — AI Foundations for Researchers | `session-1-foundations/` | 30 | 31 | 1,956 | 2,548 |
| 2 — Literature Discovery & Synthesis | `session-2-literature/` | 30 | 36 | 2,447 | 3,154 |
| 3 — Reading, Notes & Knowledge Management | `session-3-reading-km/` | 30 | 34 | 2,588 | 4,570 |
| 4 — Writing, Publishing & Integrity | `session-4-writing-integrity/` | 31 | 31 | 2,672 | 2,970 |
| 5 — Data, Code & Your AI Workflow | `session-5-data-code/` | 30 | 53 | 2,538 | 3,624 |
| 6 — AI for Multidisciplinary Research | `session-6-multidisciplinary/` | 30 | 66 | 2,536 | 3,721 |
| 7 — Discipline Deep-Dives: Specialized AI by Field | `session-7-disciplines/` | 33 | 70 | 3,350 | 6,253 |
| **Series total** | | **214** | **321** | | |
Counts re-measured 2026-09-02 after the modernization pass (slide counts unchanged; ~25 sources added series-wide). Per-type tallies (peer-reviewed / policy / primary-doc / news) live in each session's `sources.md` counts block. Every session directory contains exactly four files: `slides.html`, `sources.md`, `demo-script.md`, `handout.md`.
Also at the root: `curriculum.md` (master document — pitch, audience, format, and per-session abstract / objectives / topics / demo / takeaways), `curriculum-design.md` (the frozen design brief), `implementation-plan.md`, and `assets/` (`slide-template.html`, `handout-template.md`, `sources-template.md`).
**Source-type vocabulary across all 296 entries:** 171 `peer-reviewed` (includes preprints), 89 `primary-doc`, 29 `policy`, 5 `news`. Vendor self-reported figures are typed `primary-doc` and carry a visible provenance label on the slide that uses them.
### The seven canonical series artefacts
Named identically in every deck, handout and recap slide that references them:
1. **The Capability/Failure Map for Research Tasks** (Safe / Unreliable / Dangerous) — Session 1
2. **The 2026 Tool Landscape Taxonomy** (Tier 1 — General chatbots / Tier 2 — Research-specific tools / Tier 3 — Deep-research agents) — Session 1
3. **The Literature-Discovery Tool Matrix** — Session 2
4. **The Summary-Trust Triage** — Session 3
5. **The Publisher AI-Policy Comparison Table** — Session 4
6. **The Personal AI Workflow Canvas** — Session 5 (row 6 filled in Session 6, row 7 in Session 7; the canvas is finished at the end of the series)
7. **The Cross-Field Orientation Ladder** (Translate / Map / Read / Synthesise) — Session 6
Plus **CRIT** — Context / Role and register / Instructions and constraints / Task, then iterate — the prompting template introduced in Session 1 and reused verbatim in every later recap, and **The Field-Frontier Tracker** (Session 7).
### Verification status
#### 2026-09-02 post-modernization render pass (current)
Harness: `modernization-notes/check-deck.js` (headless Chromium, 1500px viewport — above the 1392px false-overflow floor). Per deck: overflow checked on **every element with clipping/scrolling overflow**, a geometric check that no descendant box escapes its slide (±2px), the font floor walked over **every text node** (non-exempt <18px = fail), and footer sequence asserted. The modernization pass initially left 21 overflowing slides across six decks; all were fixed by scoped spacing, prose trims (no figure, model name, era label, or citation marker lost), and reference-slide rebalancing — never by shrinking type.
| Deck | Slides | Footers sequential | Overflow | Escapes | Non-exempt <18px |
|---|---|---|---|---|---|
| Session 1 | 30 | ✅ | 0 | 0 | 0 |
| Session 2 | 30 | ✅ | 0 | 0 | 0 |
| Session 3 | 30 | ✅ | 0 | 0 | 0 |
| Session 4 | 31 | ✅ | 0 | 0 | 0 |
| Session 5 | 30 | ✅ | 0 | 0 | 0 |
| Session 6 | 30 | ✅ | 0 | 0 | 0 |
| Session 7 | 33 | ✅ | 0 | 0 | 0 |
**Print-to-PDF re-verified** on the full Session 7 deck (headless Chrome, `--print-to-pdf`): 33 pages, one slide per page, 469 KB. **Citation markers re-verified 2026-09-02**: in every session, the `[n]` markers used across slides/handout/demo exactly match the `sources.md` entry list (no orphans in either direction). **Curriculum sync re-verified 2026-09-02**: all 41 learning objectives verbatim-identical between `curriculum.md` and the decks.
#### 2026-08-08 pass (pre-modernization, for reference)
Each deck was served on its own loopback port with nothing else rendering; page identity (port + document title) was asserted immediately before every measurement; font sizes were measured by walking **every text node** and resolving its parent's computed style; overflow was checked on **every element with a clipping or scrolling overflow**, not just `.slide-inner`, plus a geometric check that no descendant's box escapes its slide.
| Deck | Slides | Footers sequential | Overflow | Min. non-exempt font | External requests |
|---|---|---|---|---|---|
| Session 1 | 30 | ✅ 1/30 … 30/30 | 0 (491 text nodes walked) | 18px | 0 |
| Session 2 | 30 | ✅ | 0 (644 nodes) | 18px | 0 |
| Session 3 | 30 | ✅ | 0 (681 nodes) | 18px | 0 |
| Session 4 | 31 | ✅ | 0 (606 nodes) | 18px | 0 |
| Session 5 | 30 | ✅ | 0 (837 nodes) | 18px | 0 |
| Session 6 | 30 | ✅ | 0 (1,180 nodes) | 18px | 0 |
| Session 7 | 33 | ✅ | 0 (1,290 nodes) | 18px | 0 |
Viewport 1500px (above the 1392px floor below which `min(1280px, 92vw)` shrinks the canvas and yields false overflow). Exempt from the 18px floor by series ruling: captions, citation superscripts, template chrome (eyebrow, footer) and the references bibliography.
**Print-to-PDF verified** on the full Session 7 deck (headless Chrome, `--print-to-pdf`): 33 pages, one slide per page, 454 KB. The `@media print` rules are identical across all seven decks.
---
## 2. Pre-delivery freshness checklist
Every deck and handout is stamped **"Landscape as of August 2026."** The facts below are the ones that move. Work through this list before the first session, and re-check the session-specific block in the week before each session. Each item names where it appears.
**New items from the 2026-09 modernization pass (check first — these move fastest)**
- [ ] **Claude Fable 5.1 / Mythos 5.1 shipped 2026-09-01** (after most deck figures were fetched). Artificial Analysis has rebased its indices (Fable 5.1 = 66; SciCode 62.0%; HLE 59.1%). Deck figures naming "Claude Fable 5" (S5 SciCode 60.2%, S6 HLE 55.5%) are correct as dated late-August measurements — decide before delivery whether to bump them to the 5.1 numbers or say the update aloud.
- [ ] **ChatGPT deep-research quotas** are stated nowhere as current (last documented figures are mid-2025); check your own account's quota before the S2 demo.
- [ ] **MLE-bench leaderboard is paused** ("developing an improved process") — re-check whether it reopened before quoting the S5 trajectory slide.
- [ ] **Onweller et al. (arXiv:2605.06635) and the Peters et al. 2026 follow-up are preprints** — check for journal publication and updated numbers (S2/S3/S6).
- [ ] **FDA AI-device list**: S7 keeps the 2026-06-16 download counts (1,524; 76.4% radiology; 0 LLM-based). FDA's foundation-model tagging plan was still future-tense on 2026-09-02; the first tagged entry changes the S7 headline.
- [ ] **Model versions in demo tools** — unchanged advice, but doubly important now: every demo script's dual-path branches were calibrated against the August-2026 generation; rehearse each demo against the model you will actually use and screenshot whichever branch fires.
**Check once, before the series opens (affects several sessions)**
- [ ] **Gemini Notebook** (renamed from NotebookLM on 16 July 2026) — product name, and the documented free-tier limits quoted in Session 3 slide 18: 50 sources per notebook, 500,000 words per source, 50 chat queries/day, 100 notebooks. Also named in Sessions 1 and 6. Source: `support.google.com/gemininotebook/answer/16269187`.
- [ ] **Model versions in the demo tools.** Every demo script tells the presenter to read the model version out of the UI on the day and say it aloud. Do this — results are version-specific and the audience will ask.
- [ ] **Vendor pricing pages**, all of which changed at least once during authoring: Elicit, Consensus, Scite, ResearchRabbit, SCiNiTO, Connected Papers, Litmaps, Julius, OpenAlex.
**Session 1**
- [ ] Slides 8, 9 and 14 rest on vendor documentation, which ages fastest in this deck.
- [ ] Tier-2 tool list on slide 20 (Elicit, Consensus, Semantic Scholar, SCiNiTO, Scite, ResearchRabbit, Gemini Notebook) — confirm all seven still exist and are still described accurately.
**Session 2**
- [ ] Seven vendor corpus/price figures in The Literature-Discovery Tool Matrix: Elicit 125M vs 138M, Consensus 200M+, Semantic Scholar 214M papers / 2.49B citations / 79M authors, Scite 280M+ vs 300M+ and 30+ publisher agreements, ResearchRabbit 310M+, SCiNiTO 500M+, OpenAlex 243M works / 48M open access. **The internal contradictions are the teaching point** — if a vendor has fixed one, update the slide and keep the lesson.
- [ ] Deep-research agent modes in ChatGPT / Gemini / Claude — names and availability change often.
**Session 3**
- [ ] Gemini Notebook limits (above) — the most perishable facts in the deck.
- [ ] Elicit, SciSpace, Paperpal, SCiNiTO help-centre wording quoted on slides 7 and 12.
- [ ] Zotero plugin ecosystem and its "plugins have full access to your Zotero and your computer" warning.
- [ ] Demo corpus: five open-access DOIs, verified against Crossref and Unpaywall on 2026-07-27 — re-resolve them.
**Session 4**
- [ ] All eight publisher policy pages quoted verbatim in The Publisher AI-Policy Comparison Table: Elsevier, Springer Nature, Wiley, Taylor & Francis, Sage, Science/AAAS, JAMA Network, ICMJE. **These are quoted byte-exactly** — if a page has changed, the quotation must change with it.
- [ ] Funder policies: NIH and NSF.
- [ ] Detector-vendor claims (the deck's point is that they are unreliable; the specific products cited may move).
**Session 5**
- [ ] **JASP 0.98's AI feature** — it was 26 days old at authoring. Confirm it still exists and still does what slide 19 says.
- [ ] **Julius** file-size limit — undocumented at authoring; check whether it has been published.
- [ ] Agentic-tool documentation quoted on slides 20–22, including the "containers expire 30 days after creation" limit.
- [ ] Demo dataset: Palmer Penguins (CC0). All demo statistics were independently recomputed from the raw CSV — they do not need re-checking, only the download link does.
**Session 6**
- [ ] **OpenAlex usage-based pricing** — "$1 of free usage every day" with a key, 2,000-char input, 50 results. Newest and most volatile item in the deck.
- [ ] `inciteful.xyz` redirect behaviour and its "100% Free to use" / 240M+ papers claims.
- [ ] Connected Papers (5 graphs/month free, USD3/mo academic), ResearchRabbit ("$0, Forever!", 50 seed articles, RR+ $10/mo → 300), Litmaps (2 maps / 100 articles free, Pro USD10/mo), Open Knowledge Maps (100 documents per map).
- [ ] Semantic Scholar Recommendations API — key requirements.
- [ ] Web of Science / Scopus / OpenAlex classification counts (4 domains / 26 fields / ~252 subfields / ~4,500 topics).
**Session 7**
- [ ] **FDA AI-Enabled Medical Device List** — 1,524 entries, content current 16 June 2026, 76.4% radiology, 0 identified as LLM-based. The counts are **author-derived from FDA's downloadable file** and the slide says so; recount from the current file.
- [ ] **Mathlib** — 283,218 theorems, 772 contributors. Live counters; they move weekly.
- [ ] **Matbench Discovery** leaderboard model count — recorded as a dated access history precisely because it moves (41 → 42 during authoring).
- [ ] Rubin Observatory alert volumes and broker count (seven full-stream brokers plus two downstream services).
- [ ] Insilico rentosertib Phase 2a/3 status (NCT07687459) and the AlphaFold 3 / ESM3 / BindCraft server terms.
- [ ] Transkribus pricing and CER figures; MAXQDA / NVivo / ATLAS.ti / Taguette AI-feature documentation.
- [ ] All eleven demo endpoints returned 200 at authoring — re-check before the poll, since two of the five demos will run live.
---
## 3. What a presenter must do before each session
**Once, before Session 1**
- [ ] **Fill in `{{Presenter Name}}`.** It appears on the title slide of all seven decks — search each `slides.html` for `{{Presenter Name}}`. By Ehsan's 2026-07-27 ruling the title slide carries the presenter's name only, with **no affiliation line**; do not add one.
- [ ] Decide how decks are distributed (HTML file, or print each to PDF from Chrome — verified working).
- [ ] Read `curriculum.md` end to end. Each deck's learning-objective slides copy that document verbatim; they are in sync as of this pass.
**Accounts and access needed for the demos**
| Session | Needed | Free tier enough? |
|---|---|---|
| 1 | Two Tier-1 chatbots (e.g. ChatGPT + Claude or Gemini), one Tier-2 tool (Consensus / Elicit / Semantic Scholar / SCiNiTO), one citation-verification tab (Google Scholar or Crossref) | Yes |
| 2 | One citation-grounded discovery tool **and** one deep-research agent mode — the agent mode usually requires a paid plan | **No** — budget for one paid agent seat |
| 3 | Gemini Notebook account; the five demo PDFs downloaded in advance | Yes (free tier limits are themselves part of the lesson) |
| 4 | Any general assistant; the target journal's live policy page open in a tab | Yes |
| 5 | A chatbot with code execution / file upload (paid on most products), or the documented no-code path; Palmer Penguins CSV downloaded | **No** for the code path; yes for the no-code variant |
| 6 | Connected Papers or Litmaps, Inciteful, OpenAlex (API key for semantic search), Semantic Scholar | Yes, but get the OpenAlex key in advance |
| 7 | PubTator3, a structure-prediction server, a text-to-CAD tool, Lean 4 Web, and a spreadsheet for Cohen's κ — five demos prepared, two run by audience poll | Yes, all free / no login for the two most likely picks |
**Per session, in the week before**
- [ ] Work the freshness block for that session (§2 above).
- [ ] **Run the demo script end to end and screenshot every step** into `fallback/01.png …`. Every script specifies this and every script has a fallback plan; live models are non-deterministic and the fallback deck is the only thing that saves a failed demo.
- [ ] Turn web search/browsing **off** where the script says so, clear chat history, disable notifications, set browser zoom to ~125%.
- [ ] Note the model version shown in each tool's UI and say it out loud at the start of the demo.
- [ ] Session 7 only: prepare the audience poll and be ready for any two of the five demos.
---
## 4. Known open items
1. **Session 7 handout is 2,983 words**, above the ~2,400-word / 2–4 page cap set by Ehsan's 2026-07-28 ruling. It was deliberately trimmed from 3,064 and left here: it carries five per-field frontier maps plus the series-wide closing checklist. Decide whether to accept the overrun or cut a field map. *(All other handouts are within cap.)*
2. **Session 4 is 31 slides and Session 7 is 33**, against a ~25–30 target. Session 4's extra slide came from splitting the policy table three ways for legibility; Session 7's from the Task 10 reference-slide ruling (§5). Both deviations are deliberate.
3. **Five sources across the series carry their marker in the handout rather than on a slide** — Session 5 `[46] [47]` and Session 7 `[6] [35] [64]`. Every one of them is cited somewhere; this is the only reason a deck's on-slide marker count can differ from its source count. Not a defect, but worth knowing if someone audits the numbers.
4. **Preprints carry point-of-use hedges** in Sessions 1, 6 and 7. Presenters should know which claims rest on preprints — Session 6 slides 11 and 17 and Session 7's social-simulation segment are the main ones. Q&A awareness, not a defect.
5. **Reference slides fit with zero spare vertical space** in Sessions 6 and 7 after the series font ruling. They are correct at the decks' fixed 1280×720 canvas. If anyone edits a reference entry to be longer, re-run the overflow check.
6. **`curriculum-design.md` was deliberately not edited.** It is the frozen input brief and still contains pre-rename wording ("NotebookLM-style"). `curriculum.md`, the live master document, has been updated.
---
## 5. Series-wide typography rulings applied in this pass
Recorded here because they now govern all seven decks:
- **Font floor.** Content text sits at **≥18px**. Accepted exemptions: captions (15px), citation superscripts (15px), template chrome — eyebrow (16px) and footer (14px) — and the references bibliography.
- **Ruling 1 — stat-card labels and every other content sub-label sit at the 18px floor.** `.statcard .statlab`, `.tool-gloss`, `.zone-gloss`, `.slide-sourcenote`, `.quotebox .attrib` and dense-table body text (`.slide--map th/td`) are now 18px in all seven decks. Where that caused a slide to overflow, the space was recovered from padding and leading — never from type.
- **Ruling 2 — the references bibliography is 15px for entries and 12px for the metadata line, in all seven decks.** Where a deck's source count will not fit at that size, the deck adds a reference slide rather than shrinking the type. Session 7 accordingly moved from three reference slides to four (32 → 33 slides); Sessions 1 and 6, which had diverged upward and downward, now match the other five.
- Map captions are 15px everywhere, matching table captions.
---
*AI for Researchers · seven sessions · Landscape as of August 2026*