Big question: Can we turn ethical concern into an auditable decision?
Research lock: 2026-08-28

Purpose of the week

This is an integration studio, not a recap lecture. Teams should leave with an evidence-backed audit structure they can reuse for the capstone. The core skill is moving repeatedly between facts, uncertainty, values, institutions, law, control, and remedy without allowing any one layer to impersonate the whole analysis.

The end-to-end audit sequence

1. Scope the actual system

Name the decision, model/service, data flows, interface, operator, institution, affected population, geographic/legal context, lifecycle stage, and downstream use. Draw what is explicitly out of scope and justify that boundary.

2. Separate claims from evidence

Create a claim ledger with columns for claim, source, source type, date, population, uncertainty, counterevidence, owner, and decision relevance. “The model is accurate” is rejected until it has a task, metric, comparator, subgroup, dataset, and time period.

3. Map stakeholders and power

List benefits, burdens, voice, control, knowledge, and remedy for direct users, data subjects, affected non-users, workers, communities, vendors, deployers, regulators, and future stakeholders. Do not equate “stakeholder consulted” with decision power.

4. Diagnose mechanisms

Use the lifecycle-harms taxonomy, a privacy information-flow map, confusion matrices, and a responsibility map. Record competing causal hypotheses rather than selecting the first plausible story.

Build a stakeholder × principle matrix: welfare/consequences, rights/duties, justice, privacy, autonomy, safety, accountability, and institutional character. Add applicable laws and standards in a separate column; law is evidence about obligations, not a substitute for ethical reasoning.

6. Pre-mortem the deployment

Imagine the system failed badly one year later. Generate failure modes across misuse, model error, dataset shift, security, privacy, human factors, organizational incentives, unequal impact, vendor failure, and absent remedy. Rank severity, likelihood, detectability, and reversibility—but preserve catastrophic low-frequency items.

7. Design controls and residual-risk decisions

For every priority risk, name preventative, detective, responsive, and remedial controls; an owner; evidence; monitoring cadence; trigger; and stop authority. Then state residual risk in plain language and identify who is authorized to accept it.

Two full workshop cases

Clearview synthesis

Teams audit the same system through five lenses:

  • Data: scraping, provenance, biometric templates, identity linkage, downstream copies.
  • Privacy: actors, purpose, consent, inference, retention, deletion.
  • Bias: false matches, subgroup evidence, database coverage, field conditions.
  • Accountability: vendor, police deployer, operator, procurement, logging, appeal.
  • Enforcement/remedy: regulator findings, cessation, deletion, verification, and jurisdiction.

The final deliverable must distinguish established facts from plausible but unproven risks.

Hiring-system reconstruction

Give teams a fictional system inspired by documented hiring controversies: résumé ranking, assessment scores, video features, historical performance labels, recruiter dashboard, and automated rejection. Ask which artifact would surface a failure earliest: datasheet, model card, impact assessment, subgroup audit, interface test, decision log, applicant notice, appeal record, or incident report. Require a minimal audit package rather than a wish list.

Audit artifacts students should produce

  1. One-page system and decision map.
  2. Claim/evidence/uncertainty ledger.
  3. Stakeholder and power map.
  4. Data lineage and purpose map.
  5. Fairness and performance table with denominators.
  6. Responsibility and authority matrix.
  7. Pre-mortem risk register.
  8. Control plan with owners and stop triggers.
  9. Affected-person notice and appeal route.
  10. Residual-risk recommendation: deploy, narrow, pilot, pause, or reject.

Framework crosswalk

The SMACTR paper's audit stages—scoping, mapping, artifact collection, testing, reflection, and post-audit—can be crosswalked to NIST's Govern, Map, Measure, and Manage functions. The value of the crosswalk is not to declare equivalence; it lets students see which activities are missing and whether evidence survives organizational handoffs.

Canada's Algorithmic Impact Assessment is useful as a concrete public-sector instrument. As of the research lock, the current tool describes 65 risk questions and 41 mitigation questions and is mandatory under the federal Directive for covered systems. Have students critique what a questionnaire can reveal, what it can incentivize teams to minimize, and what independent validation remains necessary.

Seminar run-of-show

  • 0:00–0:25: individual audit of a deliberately incomplete system card.
  • 0:25–1:00: teams build scope and claim ledgers; facilitator injects conflicting evidence.
  • 1:00–1:30: stakeholder matrix and pre-mortem.
  • 1:30–1:40: break with unresolved risk visible.
  • 1:40–2:15: testing and control design.
  • 2:15–2:40: adversarial review by another team.
  • 2:40–3:00: revise recommendation and state residual uncertainty.

Visual evidence plan

VisualCapture targetTeaching useGuardrail
NIST RMF corepublic/courses/mai-105/evidence/nist-ai-rmf.jpgPlace each audit artifact under Govern, Map, Measure, or Manage.The functions are not a linear checklist.
Lifecycle harmspublic/courses/mai-105/evidence/lifecycle-harms.jpgSeed the pre-mortem beyond model error.Add security, organizational, and remedy failures.
Clearview findingspublic/courses/mai-105/evidence/clearview-opc.jpgBuild a multi-domain evidence board from one primary record.Separate regulator finding from broader claims.
Canadian AIAofficial AIA pageCapture risk-level logic and question structure.Covered federal use is not every Canadian deployment.
NIST Playbookofficial playbookAssign a concrete action and evidence artifact to each team.Suggested actions require contextual tailoring.

Reading and citation ledger

  1. Raji et al., “Closing the AI Accountability Gap” / SMACTR.
  2. NIST AI RMF 1.0 and NIST AI RMF Playbook.
  3. Canada, Algorithmic Impact Assessment and Directive on Automated Decision-Making.
  4. Office of the Privacy Commissioner of Canada, Clearview AI findings.
  5. Suresh & Guttag, lifecycle sources of harm.
  6. Mitchell et al., Model Cards and Gebru et al., Datasheets.

Quality gate for capstone drafts

  • Every important claim has a dated source and calibrated certainty.
  • The system boundary includes institution and remedy.
  • Metrics connect to people and decisions.
  • Controls have owners, evidence, triggers, and authority.
  • Counterevidence changed at least one recommendation.
  • The final recommendation names residual risk and who bears it.