Big question: Which recommendation survives adversarial evidence, changed assumptions, and changed law?
Research lock: 2026-08-28
Purpose of the week
The capstone should demonstrate a durable analytical practice, not recall of thirteen weeks of vocabulary. Students must scope a real socio-technical system, support empirical claims, expose normative choices, map responsibility, test counterarguments, and propose controls and remedies that an institution could actually operate.
The capstone argument
A strong audit can be summarized as:
For this decision and population, the evidence supports these benefits and risks with this uncertainty. Because these stakeholders, rights, and distributional effects matter, and because these actors have these duties and controls, we recommend deploy / narrow / pilot / pause / reject, subject to these tests, owners, triggers, remedies, and review dates.
If any bold component is missing, the recommendation is probably under-specified.
Required evidence package
- System boundary: model/service, data, interface, workflow, institution, affected people, downstream decisions, and excluded scope.
- Claim ledger: primary record, date, population, method, uncertainty, contradiction, and decision relevance.
- Data and purpose lineage: source, consent/legal basis, transformations, labels, proxies, retention, derivatives, and deletion.
- Performance/fairness evidence: comparator, threshold, denominators, subgroups, uncertainty, field conditions, and drift.
- Privacy flow: actors, attributes, purpose, recipients, inference, retention, security, and contestation.
- Responsibility map: causal contribution, knowledge, control, duty, evidence, authority, and remedy.
- Legal/governance map: applicable instrument, actor, obligation, effective date, regulator, and remedy.
- Pre-mortem and red-team results: priority failure modes, counterevidence, and changed recommendations.
- Control plan: preventative, detective, responsive, and remedial controls with owners and stop triggers.
- Residual-risk decision: who accepts it, who bears it, when it is reviewed, and what would reverse the decision.
Capstone defence format
- 4 minutes: system, decision, affected population, and recommendation.
- 4 minutes: strongest evidence and greatest uncertainty.
- 4 minutes: stakeholder/ethical conflict and legal duties.
- 4 minutes: controls, monitoring, remedy, and stop authority.
- 6 minutes: adversarial panel.
- 3 minutes: revised recommendation after the challenge.
The revision is graded. Refusing to update in the face of relevant evidence is not confidence.
Adversarial review prompts
Evidence
- Which claim depends on a vendor or advocacy source?
- Which result fails to transfer to this population or date?
- What denominator, comparator, or counterfactual is missing?
- What observation would falsify the core benefit claim?
Values and distribution
- Which affected group did not participate in defining success?
- Which right is treated as a tradeable average benefit?
- Who receives gains, and who receives error, delay, surveillance, or appeal burden?
System and institution
- Which failure could occur even if the model performs exactly as specified?
- Is the “human in the loop” informed, resourced, and authorized?
- What incentive causes a written control to be ignored?
- What evidence will exist after an incident?
Remedy and future change
- Can an affected person learn about, challenge, and change the decision?
- What happens after model, data, law, vendor, or use changes?
- Who can stop deployment without executive permission?
Future scenarios as disciplined uncertainty
AI 2027: audit a forecast, not a vibe
Scenario work can expose strategic assumptions, but narrative detail can feel like evidence. Students extract claims about compute, algorithmic progress, automation of AI research, organizational behaviour, geopolitics, and timelines. For each, record a measurable indicator, source, base rate, alternative, and update condition. The goal is not to endorse or dismiss the scenario; it is to make it forecastable.
Frontier safety frameworks: voluntary governance under moving capability
Compare three provider frameworks:
- Anthropic's Responsible Scaling Policy uses capability thresholds, safeguards, risk reports, and a public change log.
- OpenAI's Preparedness Framework tracks severe-harm capability categories, capabilities reports, safeguards reports, and governance review.
- Google DeepMind's Frontier Safety Framework uses tracked/critical capability levels, evaluations, mitigations, and external input.
Ask which thresholds are measurable, who validates them, what is public, what happens at a threshold, whether leaders can override, and how policy changes are governed. These are primary statements of provider process, not independent proof of safety.
AI welfare: allocate attention under moral uncertainty
The AI-welfare literature asks whether future systems could merit moral consideration and how uncertainty should influence action. Students should separate current documented human/animal/environmental harms from speculative future moral patients, then examine low-cost precautions, opportunity costs, indicators of morally relevant properties, and institutional capture. Moral uncertainty does not require either certainty or neglect.
Synthesis activities
- Week 1 belief update: students revisit their original narrative camp and state one claim strengthened, one weakened, and one evidence gap.
- Red-team carousel: every team reviews another's evidence, values, system boundary, and remedy separately.
- Policy shock: after recommendations, reveal a law change, vendor acquisition, population shift, or incident. Teams revise controls and ownership.
- One-year audit plan: specify monthly/quarterly evidence, affected-person feedback, incident review, external evaluation, and retirement condition.
Visual evidence plan
| Visual | Capture target | Teaching use | Guardrail |
|---|---|---|---|
| Audit lifecycle | SMACTR paper | Build a loop from scoping through post-audit monitoring. | An internal audit needs independence and remedy to have force. |
| NIST RMF core | public/courses/mai-105/evidence/nist-ai-rmf.jpg | Place capstone artifacts under Govern, Map, Measure, Manage. | Not a certification or exhaustive checklist. |
| Current incident evidence | OECD AI Incidents and Hazards Monitor | Find an incident, then trace it to primary records. | Automated monitoring is a lead source, not a final authority. |
| Anthropic RSP history | current RSP page | Capture version history and one threshold/action relationship. | Provider-authored and voluntary; note redactions and revisions. |
| OpenAI Preparedness | Framework update | Compare capability reports, safeguards reports, and decision governance. | Primary description, not independent evaluation. |
| DeepMind FSF | frontier-safety hub | Compare tracked/critical levels and mitigation sequence. | Primary description, not independent evaluation. |
| Scenario assumptions | AI 2027 | Screenshot a dated claim and attach indicators/alternatives. | A detailed scenario is not a probability estimate by itself. |
Reading and citation ledger
Audit and current evidence
- Raji et al., “Closing the AI Accountability Gap”.
- NIST AI RMF and AIRC implementation resources.
- International AI Safety Report 2026.
- OECD AI Incidents and Hazards Monitor.
Forecasting and frontier governance
- AI Futures Project, AI 2027 — scenario to audit, not assigned consensus.
- Anthropic, current Responsible Scaling Policy and change log.
- OpenAI, Preparedness Framework update and Framework PDF.
- Google DeepMind, Frontier Safety Framework hub and 2025/2026 update.
- Long et al., “Taking AI Welfare Seriously”.
Assessment rubric
| Dimension | Strong evidence |
|---|---|
| Scope | Decision, institution, affected people, and downstream effects are explicit. |
| Empirics | Primary sources, methods, denominators, uncertainty, and counterevidence are visible. |
| Ethics | Normative conflicts are defended rather than hidden behind metrics. |
| Governance | Current rules are mapped to actors, dates, evidence, enforcement, and remedy. |
| Feasibility | Controls have owners, resources, triggers, and verification. |
| Contestability | Affected people can discover, understand, challenge, and obtain correction. |
| Adaptation | Recommendation changes appropriately under new evidence or context. |
Final quality gate
- No unsupported superlatives or universal claims.
- No complaint described as a judgment or authorization described as outcome evidence.
- No fairness metric without stakeholder harm and denominators.
- No explanation without audience, purpose, and validation.
- No “human oversight” without information, time, competence, and authority.
- No control without an owner, artifact, trigger, and remedy.
- No future scenario without measurable assumptions and update conditions.