Big question: Can we turn ethical concern into an auditable decision?
Research lock: 2026-08-28
Purpose of the week
This is an integration studio, not a recap lecture. Teams should leave with an evidence-backed audit structure they can reuse for the capstone. The core skill is moving repeatedly between facts, uncertainty, values, institutions, law, control, and remedy without allowing any one layer to impersonate the whole analysis.
The end-to-end audit sequence
1. Scope the actual system
Name the decision, model/service, data flows, interface, operator, institution, affected population, geographic/legal context, lifecycle stage, and downstream use. Draw what is explicitly out of scope and justify that boundary.
2. Separate claims from evidence
Create a claim ledger with columns for claim, source, source type, date, population, uncertainty, counterevidence, owner, and decision relevance. “The model is accurate” is rejected until it has a task, metric, comparator, subgroup, dataset, and time period.
3. Map stakeholders and power
List benefits, burdens, voice, control, knowledge, and remedy for direct users, data subjects, affected non-users, workers, communities, vendors, deployers, regulators, and future stakeholders. Do not equate “stakeholder consulted” with decision power.
4. Diagnose mechanisms
Use the lifecycle-harms taxonomy, a privacy information-flow map, confusion matrices, and a responsibility map. Record competing causal hypotheses rather than selecting the first plausible story.
5. Apply ethical and legal lenses
Build a stakeholder × principle matrix: welfare/consequences, rights/duties, justice, privacy, autonomy, safety, accountability, and institutional character. Add applicable laws and standards in a separate column; law is evidence about obligations, not a substitute for ethical reasoning.
6. Pre-mortem the deployment
Imagine the system failed badly one year later. Generate failure modes across misuse, model error, dataset shift, security, privacy, human factors, organizational incentives, unequal impact, vendor failure, and absent remedy. Rank severity, likelihood, detectability, and reversibility—but preserve catastrophic low-frequency items.
7. Design controls and residual-risk decisions
For every priority risk, name preventative, detective, responsive, and remedial controls; an owner; evidence; monitoring cadence; trigger; and stop authority. Then state residual risk in plain language and identify who is authorized to accept it.
Two full workshop cases
Clearview synthesis
Teams audit the same system through five lenses:
- Data: scraping, provenance, biometric templates, identity linkage, downstream copies.
- Privacy: actors, purpose, consent, inference, retention, deletion.
- Bias: false matches, subgroup evidence, database coverage, field conditions.
- Accountability: vendor, police deployer, operator, procurement, logging, appeal.
- Enforcement/remedy: regulator findings, cessation, deletion, verification, and jurisdiction.
The final deliverable must distinguish established facts from plausible but unproven risks.
Hiring-system reconstruction
Give teams a fictional system inspired by documented hiring controversies: résumé ranking, assessment scores, video features, historical performance labels, recruiter dashboard, and automated rejection. Ask which artifact would surface a failure earliest: datasheet, model card, impact assessment, subgroup audit, interface test, decision log, applicant notice, appeal record, or incident report. Require a minimal audit package rather than a wish list.
Audit artifacts students should produce
- One-page system and decision map.
- Claim/evidence/uncertainty ledger.
- Stakeholder and power map.
- Data lineage and purpose map.
- Fairness and performance table with denominators.
- Responsibility and authority matrix.
- Pre-mortem risk register.
- Control plan with owners and stop triggers.
- Affected-person notice and appeal route.
- Residual-risk recommendation: deploy, narrow, pilot, pause, or reject.
Framework crosswalk
The SMACTR paper's audit stages—scoping, mapping, artifact collection, testing, reflection, and post-audit—can be crosswalked to NIST's Govern, Map, Measure, and Manage functions. The value of the crosswalk is not to declare equivalence; it lets students see which activities are missing and whether evidence survives organizational handoffs.
Canada's Algorithmic Impact Assessment is useful as a concrete public-sector instrument. As of the research lock, the current tool describes 65 risk questions and 41 mitigation questions and is mandatory under the federal Directive for covered systems. Have students critique what a questionnaire can reveal, what it can incentivize teams to minimize, and what independent validation remains necessary.
Seminar run-of-show
- 0:00–0:25: individual audit of a deliberately incomplete system card.
- 0:25–1:00: teams build scope and claim ledgers; facilitator injects conflicting evidence.
- 1:00–1:30: stakeholder matrix and pre-mortem.
- 1:30–1:40: break with unresolved risk visible.
- 1:40–2:15: testing and control design.
- 2:15–2:40: adversarial review by another team.
- 2:40–3:00: revise recommendation and state residual uncertainty.
Visual evidence plan
| Visual | Capture target | Teaching use | Guardrail |
|---|---|---|---|
| NIST RMF core | public/courses/mai-105/evidence/nist-ai-rmf.jpg | Place each audit artifact under Govern, Map, Measure, or Manage. | The functions are not a linear checklist. |
| Lifecycle harms | public/courses/mai-105/evidence/lifecycle-harms.jpg | Seed the pre-mortem beyond model error. | Add security, organizational, and remedy failures. |
| Clearview findings | public/courses/mai-105/evidence/clearview-opc.jpg | Build a multi-domain evidence board from one primary record. | Separate regulator finding from broader claims. |
| Canadian AIA | official AIA page | Capture risk-level logic and question structure. | Covered federal use is not every Canadian deployment. |
| NIST Playbook | official playbook | Assign a concrete action and evidence artifact to each team. | Suggested actions require contextual tailoring. |
Reading and citation ledger
- Raji et al., “Closing the AI Accountability Gap” / SMACTR.
- NIST AI RMF 1.0 and NIST AI RMF Playbook.
- Canada, Algorithmic Impact Assessment and Directive on Automated Decision-Making.
- Office of the Privacy Commissioner of Canada, Clearview AI findings.
- Suresh & Guttag, lifecycle sources of harm.
- Mitchell et al., Model Cards and Gebru et al., Datasheets.
Quality gate for capstone drafts
- Every important claim has a dated source and calibrated certainty.
- The system boundary includes institution and remedy.
- Metrics connect to people and decisions.
- Controls have owners, evidence, triggers, and authority.
- Counterevidence changed at least one recommendation.
- The final recommendation names residual risk and who bears it.