Training data is now a governance surface: provenance, consent, licensing, deletion, and model remediation travel together.
Data Ethics & Bias
Curating datasets and hidden proxies
If the model is a mirror of the data, who curated the reflection?
Students trace harm before modelling begins: problem formulation, consent, categories, labels, proxy targets, collection, documentation, downstream copies, and feedback loops.
Join the live room
Before class
Bring a dataset or data source you have used. Sketch its origin, labels, missing populations, intended use, and one downstream use its creators may not have anticipated.
After class
Draft a one-page datasheet for the dataset you brought. Mark one field as a hard governance gate and explain who has authority to enforce it.
The promise
By the end of this room…
- 01Locate distinct sources of harm across the machine-learning lifecycle.
- 02Distinguish a construct from its operational proxy and identify where the relationship breaks.
- 03Interrogate dataset categories, annotation labour, consent, provenance, and downstream reuse.
- 04Build a datasheet that functions as a governance boundary rather than a documentation ritual.
- 05Propose a remedy proportionate to both the original data practice and its model descendants.
Why this week now
Signals, not scene-setting.
Synthetic and generated data do not remove upstream judgments; they can harden them while making lineage harder to see.
The operational question has shifted from whether a dataset is biased to who can inspect, contest, repair, and retire it.
Run of show
Provoke → frame → work → argue → synthesize.
- 0:00provocation
Delete the data—or delete the model?
Students vote on whether a model trained on data collected without valid consent may remain in service after the source data is deleted.
Live activity · data deletion vote - 0:15frame
The bias cascade
Problem formulation → collection → measurement → representation → aggregation → evaluation → deployment → feedback. Each stage creates a different repair obligation.
- 0:45discussion
Construct, proxy, missing world
Pairs unpack a familiar target—engagement, risk, employability, health need—and name what the proxy rewards, excludes, and teaches downstream systems.
Live activity · proxy audit - 1:10frame
Datasets are institutions
Categories have histories; labels contain labour; consent is contextual; documentation and versioning decide whether repair can propagate.
- 1:35break
Break
Ten minutes. Leave the room's unresolved question visible.
- 1:45forensics
Dataset archaeology
Groups examine Tay and ImageNet through four lenses: provenance, taxonomy, labour, and deployment. Each group builds a causal chain rather than a list of bad outcomes.
Deliverable · A pipeline map with one high-leverage intervention and one unresolved trade-off.Live activity · dataset archaeology - 2:25controversy
Remediation tribunal
The room decides what repair requires when data was collected illegitimately but a socially useful model already exists.
Live activity · remedy ranking - 2:50synthesis
A datasheet that can stop deployment
Each group writes the one datasheet field that should trigger review, refusal, or retirement rather than merely record a caveat.
Live activity · week 2 exit
Case room
Evidence before opinion.
Microsoft Tay
Which failure was data, which was interaction design, and which was operations?
ImageNet person categories
What makes a taxonomy an exercise of institutional power?
Everalbum
When must remedy reach models derived from improperly retained data?
Hidden proxy pipeline
Can accurate prediction still be invalid because the target is wrong?
Reading stack
Read the tension, not the bibliography.
- 01CoreA Framework for Understanding Sources of Harm ↗
Harini Suresh & John Guttag
- 02CoreDatasheets for Datasets ↗
Timnit Gebru et al.
- 03CurrentExcavating AI ↗
Kate Crawford & Trevor Paglen
- 04CurrentLearning from Tay's introduction ↗
Microsoft
Evidence ledger
Every case has a receipt.
6 primary, scholarly, or first-party sources
A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle
Separates historical, representation, measurement, aggregation, evaluation, and deployment harms.
Research framework ↗Datasheets for Datasets
A structured record for dataset motivation, composition, collection, uses, maintenance, and limitations.
Documentation framework ↗
Learning from Tay's introduction
A compact case about adversarial interaction, filters, monitoring, escalation, and shutdown authority.
First-party postmortem ↗
Excavating AI
Makes ImageNet's person categories and their institutional genealogy visible and contestable.
Critical visual investigation ↗
ImageNet update on person categories
A first-party account of removing unsafe and offensive person categories from a major dataset.
Dataset remediation statement ↗FTC finalizes Everalbum facial-recognition settlement
A consent-and-remedy case where models developed from improperly retained data also had to be deleted.
Regulatory settlement ↗