// course · read in orderest. 2026 · no ads · anonymous stats
W02Full seminar3 hours

Data Ethics & Bias

Curating datasets and hidden proxies

If the model is a mirror of the data, who curated the reflection?

Students trace harm before modelling begins: problem formulation, consent, categories, labels, proxy targets, collection, documentation, downstream copies, and feedback loops.

Join the live room
A data-to-decision pipeline showing where choices and harms enter
Visual field note · Week 2 · use this diagram to keep the model inside its surrounding system.

Before class

Bring a dataset or data source you have used. Sketch its origin, labels, missing populations, intended use, and one downstream use its creators may not have anticipated.

After class

Draft a one-page datasheet for the dataset you brought. Mark one field as a hard governance gate and explain who has authority to enforce it.

The promise

By the end of this room…

  1. 01Locate distinct sources of harm across the machine-learning lifecycle.
  2. 02Distinguish a construct from its operational proxy and identify where the relationship breaks.
  3. 03Interrogate dataset categories, annotation labour, consent, provenance, and downstream reuse.
  4. 04Build a datasheet that functions as a governance boundary rather than a documentation ritual.
  5. 05Propose a remedy proportionate to both the original data practice and its model descendants.

Why this week now

Signals, not scene-setting.

01

Training data is now a governance surface: provenance, consent, licensing, deletion, and model remediation travel together.

02

Synthetic and generated data do not remove upstream judgments; they can harden them while making lineage harder to see.

03

The operational question has shifted from whether a dataset is biased to who can inspect, contest, repair, and retire it.

Run of show

Provoke → frame → work → argue → synthesize.

Open student room ↗
  1. 0:00
    provocation

    Delete the data—or delete the model?

    Students vote on whether a model trained on data collected without valid consent may remain in service after the source data is deleted.

    Live activity · data deletion vote
  2. 0:15
    frame

    The bias cascade

    Problem formulation → collection → measurement → representation → aggregation → evaluation → deployment → feedback. Each stage creates a different repair obligation.

  3. 0:45
    discussion

    Construct, proxy, missing world

    Pairs unpack a familiar target—engagement, risk, employability, health need—and name what the proxy rewards, excludes, and teaches downstream systems.

    Live activity · proxy audit
  4. 1:10
    frame

    Datasets are institutions

    Categories have histories; labels contain labour; consent is contextual; documentation and versioning decide whether repair can propagate.

  5. 1:35
    break

    Break

    Ten minutes. Leave the room's unresolved question visible.

  6. 1:45
    forensics

    Dataset archaeology

    Groups examine Tay and ImageNet through four lenses: provenance, taxonomy, labour, and deployment. Each group builds a causal chain rather than a list of bad outcomes.

    Deliverable · A pipeline map with one high-leverage intervention and one unresolved trade-off.Live activity · dataset archaeology
  7. 2:25
    controversy

    Remediation tribunal

    The room decides what repair requires when data was collected illegitimately but a socially useful model already exists.

    Live activity · remedy ranking
  8. 2:50
    synthesis

    A datasheet that can stop deployment

    Each group writes the one datasheet field that should trigger review, refusal, or retirement rather than merely record a caveat.

    Live activity · week 2 exit

Case room

Evidence before opinion.

Anchor

Microsoft Tay

Which failure was data, which was interaction design, and which was operations?

Microsoft
Counterpoint

Hidden proxy pipeline

Can accurate prediction still be invalid because the target is wrong?

EAAMO / arXiv

Reading stack

Read the tension, not the bibliography.

  1. 01
    CoreA Framework for Understanding Sources of Harm

    Harini Suresh & John Guttag

  2. 02
    CoreDatasheets for Datasets

    Timnit Gebru et al.

  3. 03
    CurrentExcavating AI

    Kate Crawford & Trevor Paglen

  4. 04
    CurrentLearning from Tay's introduction

    Microsoft

Evidence ledger

Every case has a receipt.

6 primary, scholarly, or first-party sources