Frontier-model interpretability now investigates internal computation, but useful traces remain partial and contestable.
Explainable AI
Transparency, trust, and the black box
An explanation for whom, faithful to what, and useful for which action?
Students compare interpretable models, post-hoc explanations, counterfactuals, procedural explanations, mechanistic interpretability, and chain-of-thought monitoring by audience and purpose.
Join the live roomBefore class
Bring one explanation produced by an AI product. Identify its audience, intended action, and untested faithfulness claim.
After class
Redesign the explanation for an affected person who needs to contest the outcome.
The promise
By the end of this room…
- 01Distinguish interpretability, explainability, transparency, and contestability.
- 02Evaluate local and global explanation faithfulness.
- 03Select an explanation for a developer, regulator, user, or affected person.
- 04Know when an inherently interpretable model should be preferred.
Why this week now
Signals, not scene-setting.
An explanation can be persuasive without being faithful and detailed without being actionable.
The relevant standard for an affected person is often whether they can correct, challenge, and obtain remedy.
Run of show
Provoke → frame → work → argue → synthesize.
- 0:00provocation
Commit before the concepts
A black-box model is 1% more accurate than an interpretable model in a high-stakes setting. Which should be deployed?
Live activity · week 10 opening - 0:15frame
Explanation is an audience–purpose contract
Frontier-model interpretability now investigates internal computation, but useful traces remain partial and contestable.
- 0:45discussion
Reading tension
Student leaders present the assigned readings as a clash of defensible positions, then moderate questions that expose the hidden assumptions.
- 1:10frame
Faithfulness and the interpretability frontier
An explanation can be persuasive without being faithful and detailed without being actionable.
- 1:35break
Break
Ten minutes. Leave the room's unresolved question visible.
- 1:45forensics
Wolf, husky, and the shortcut hypothesis
Role-based groups work the anchor case through technical, legal, stakeholder, and normative lenses.
Deliverable · A two-minute finding with evidence, uncertainty, and an actionable remedy.Live activity · week 10 forensics - 2:25controversy
Stakeholder explanation design studio
Assigned positions, side-switch, and a joint recommendation that names the value or stakeholder it leaves exposed.
Live activity · week 10 controversy - 2:50synthesis
Re-vote and leave a trace
Repeat the opening vote, inspect what moved, and submit the strongest argument you still reject.
Live activity · week 10 exit
Case room
Evidence before opinion.
Wolf versus husky
What did the local explanation establish—and what did it not?
Circuit tracing
When does a computational graph become evidence about a model's behaviour?
Reading stack
Read the tension, not the bibliography.
- 01CoreStop explaining black box models ↗
Cynthia Rudin
- 02CoreWhy Should I Trust You? ↗
Ribeiro, Singh & Guestrin
- 03CurrentCircuit Tracing ↗
Anthropic
Evidence ledger
Every case has a receipt.
3 primary, scholarly, or first-party sources
Stop explaining black box machine learning models for high stakes decisions
Challenges the assumed universal trade-off between accuracy and interpretability.
Peer-reviewed perspective ↗Why Should I Trust You? Explaining the Predictions of Any Classifier
Introduces LIME and the wolf-versus-husky shortcut example while motivating faithfulness questions.
Research paper ↗Circuit Tracing: Revealing Computational Graphs in Language Models
A current example of tracing internal model computation rather than only explaining outputs after the fact.
Mechanistic interpretability research ↗