Big question: What evidence is enough when an error changes care?
Research lock: 2026-08-28
Why this week matters
Healthcare makes abstract trade-offs concrete. A model may improve average discrimination yet fail under dataset shift, change clinician workflow, redistribute scarce resources, or produce automation bias. Authorization is not proof of benefit in every setting, and a retrospective benchmark is not a clinical outcome. Students need to connect technical evidence to intended use, patient population, workflow, regulation, monitoring, and remedy.
Deeper teaching spine
1. Start with the clinical pathway
Map symptoms or screening invitation → data acquisition → model → result → clinician interpretation → action → patient outcome → follow-up. State whether the tool screens, diagnoses, triages, predicts, recommends, drafts, or automates. The same performance number can have different consequences at different points.
2. Build an evidence ladder
- Technical validity on a held-out dataset.
- External validation across sites and populations.
- Prospective workflow evaluation.
- Comparison with current standard of care.
- Patient outcomes, not only model metrics.
- Subgroup safety and access effects.
- Post-deployment monitoring under change.
Ask where each study sits. Do not allow a high AUROC to become a claim of lives saved without the intermediate evidence.
3. Prevalence changes meaning
Sensitivity and specificity do not tell a clinician how many alerts will be true in a given population. Use Bayes' rule or a natural-frequency table. A test with 90% sensitivity and 90% specificity deployed where prevalence is 1% produces roughly 9 true positives and 99 false positives per 1,000 people. Workload and harm follow from context.
4. Dataset shift is normal, not exceptional
Sites differ in equipment, demographics, coding, treatment, disease prevalence, referral rules, and missingness. Practices change after deployment. Require an intended-population statement, site acceptance test, drift indicators, outcome monitoring, and a stop/revalidation trigger.
5. Human oversight must fit clinical work
Measure alert volume, time, information, trust calibration, override, escalation, and accountability. A clinician may technically remain responsible while lacking access to model limitations or the time to investigate. Conversely, dismissing useful decision support can also harm patients. Design calibrated reliance, not reflexive acceptance or rejection.
6. Regulation is lifecycle governance
The FDA's AI-enabled medical-device list identifies authorized devices based largely on AI-related language in authorization summaries; it is not a quality ranking. Current digital-health guidance emphasizes total product lifecycle, real-world performance, cybersecurity, transparency, and predetermined change-control plans. Students should distinguish a static clearance from the governed evolution of a software function.
Case-study dossier
Obermeyer: resource allocation through a distorted target
Revisit the cost-as-need proxy in a health setting. The intervention should change the target or allocation logic, not simply remove race. Ask which clinical variables better represent need and whether the healthcare system itself produces their missingness.
Epic Sepsis Model: external validation before reliance
Wong et al. externally evaluated a widely implemented proprietary sepsis model and reported substantially lower discrimination and sensitivity than vendor information had suggested in their setting, along with a high alert burden. Use the paper to ask what hospitals should demand before procurement: local validation, threshold data, prospective testing, subgroup analysis, update notice, and access to failure logs.
Diabetic-retinopathy screening in Thailand: accuracy meets workflow
Google researchers' human-centred clinic study found that image-quality thresholds, network conditions, clinic workflows, and patient expectations shaped deployment. A later prospective multisite study reported specialist-comparable performance for vision-threatening disease while emphasizing socio-environmental factors. Pairing the studies demonstrates why field implementation evidence complements accuracy.
AI device updates: when the product changes
Students design a predetermined change-control plan for a screening model: what changes are anticipated, how data represent intended populations, what tests run before and after, what users are told, what monitoring detects degradation, and when a new regulatory submission or rollback is required.
Seminar activities
- Natural-frequency clinic: translate sensitivity, specificity, and prevalence into 1,000-patient outcomes; allocate follow-up capacity.
- Evidence tribunal: vendor, hospital, clinician, patient, and regulator decide whether evidence supports pilot, deployment, or rejection.
- Silent shift: facilitator changes scanner type, population, clinical protocol, and prevalence. Teams identify which evidence no longer transfers.
- Override review: analyze ten clinician overrides. Decide whether they reveal model error, workflow mismatch, inappropriate trust, or documentation gaps.
Visual evidence plan
| Visual | Capture target | Teaching use | Guardrail |
|---|---|---|---|
| Obermeyer causal receipt | public/courses/mai-105/evidence/obermeyer-health.jpg | Show how spending entered as a proxy for need. | Avoid biological interpretations of race. |
| FDA device list | FDA AI-enabled devices | Capture intended-use/category fields from two contrasting devices. | Listing means authorization under applicable requirements, not universal superiority. |
| Epic sepsis validation | JAMA Internal Medicine article | Crop study design and principal validation results. | One external site does not establish every site's performance. |
| Thailand clinic workflow | Google Research CHI study | Capture the abstract's data-quality/workflow tensions. | This is a first-party research team; teach methods and limitations. |
| Prospective screening study | Google Research / Lancet Digital Health | Compare accuracy, sensitivity, specificity, PPV, and NPV. | Metrics belong to the studied programme and prevalence. |
| Change-control plan | FDA principles | Highlight evidence before/after a planned change and monitoring. | Guidance status and jurisdiction must be labelled. |
Reading and citation ledger
- WHO, Ethics and governance of artificial intelligence for health.
- FDA, AI-enabled medical-device list.
- FDA, digital-health guidance index — check current final/draft status before class.
- FDA/Health Canada/MHRA, predetermined change-control principles.
- Wong et al., external validation of a proprietary sepsis model.
- Beede et al., human-centred evaluation of diabetic-retinopathy AI in Thailand.
- Raumviboonsuk et al., prospective multisite screening study.
- Obermeyer et al., racial bias in population-health management.
Watch list
- Check FDA list and guidance status immediately before teaching.
- Never quote performance without intended use, comparator, population, prevalence, and threshold.
- Separate authorization, clinical adoption, reimbursement, benefit, and safety; they are different claims.