// course · read in orderest. 2026 · no ads · anonymous stats
W10Course shell3 hours

Explainable AI

Transparency, trust, and the black box

An explanation for whom, faithful to what, and useful for which action?

Students compare interpretable models, post-hoc explanations, counterfactuals, procedural explanations, mechanistic interpretability, and chain-of-thought monitoring by audience and purpose.

Join the live room

Before class

Bring one explanation produced by an AI product. Identify its audience, intended action, and untested faithfulness claim.

After class

Redesign the explanation for an affected person who needs to contest the outcome.

The promise

By the end of this room…

  1. 01Distinguish interpretability, explainability, transparency, and contestability.
  2. 02Evaluate local and global explanation faithfulness.
  3. 03Select an explanation for a developer, regulator, user, or affected person.
  4. 04Know when an inherently interpretable model should be preferred.

Why this week now

Signals, not scene-setting.

01

Frontier-model interpretability now investigates internal computation, but useful traces remain partial and contestable.

02

An explanation can be persuasive without being faithful and detailed without being actionable.

03

The relevant standard for an affected person is often whether they can correct, challenge, and obtain remedy.

Run of show

Provoke → frame → work → argue → synthesize.

Open student room ↗
  1. 0:00
    provocation

    Commit before the concepts

    A black-box model is 1% more accurate than an interpretable model in a high-stakes setting. Which should be deployed?

    Live activity · week 10 opening
  2. 0:15
    frame

    Explanation is an audience–purpose contract

    Frontier-model interpretability now investigates internal computation, but useful traces remain partial and contestable.

  3. 0:45
    discussion

    Reading tension

    Student leaders present the assigned readings as a clash of defensible positions, then moderate questions that expose the hidden assumptions.

  4. 1:10
    frame

    Faithfulness and the interpretability frontier

    An explanation can be persuasive without being faithful and detailed without being actionable.

  5. 1:35
    break

    Break

    Ten minutes. Leave the room's unresolved question visible.

  6. 1:45
    forensics

    Wolf, husky, and the shortcut hypothesis

    Role-based groups work the anchor case through technical, legal, stakeholder, and normative lenses.

    Deliverable · A two-minute finding with evidence, uncertainty, and an actionable remedy.Live activity · week 10 forensics
  7. 2:25
    controversy

    Stakeholder explanation design studio

    Assigned positions, side-switch, and a joint recommendation that names the value or stakeholder it leaves exposed.

    Live activity · week 10 controversy
  8. 2:50
    synthesis

    Re-vote and leave a trace

    Repeat the opening vote, inspect what moved, and submit the strongest argument you still reject.

    Live activity · week 10 exit

Case room

Evidence before opinion.

Canonical

Wolf versus husky

What did the local explanation establish—and what did it not?

KDD / arXiv
Current

Circuit tracing

When does a computational graph become evidence about a model's behaviour?

Anthropic

Reading stack

Read the tension, not the bibliography.

  1. 01
    CoreStop explaining black box models

    Cynthia Rudin

  2. 02
    CoreWhy Should I Trust You?

    Ribeiro, Singh & Guestrin

  3. 03
    CurrentCircuit Tracing

    Anthropic

Evidence ledger

Every case has a receipt.

3 primary, scholarly, or first-party sources