Thu, Sep 17
Topic 2. How Computers Work
Learning From Data: Training Sets, Errors, and Bias
Please complete the tasks and readings listed below before class on Thu, Sep 17.
Required Readings
Optional Readings
Reading
Google. Teachable Machine [Interactive]. Link
Try object classification; test unfamiliar backgrounds and examples.
Reading
U.S. Food and Drug Administration. “FDA Proposes Updated Recommendations to Help Improve Performance of Pulse Oximeters Across Skin Tones.” FDA, 6 Jan. 2025. Link
Measurement bias; consider how errors can enter a system before machine learning begins.
Reading
Koenecke, Allison, et al. “Automated Speech Recognition Less Accurate for Blacks.” Stanford Report, 23 Mar. 2020. Link
Interactive examples show substantially higher speech-recognition error rates for Black speakers.
Reading
Stanford Computational Policy Lab. “The Race Gap in Speech Recognition Technology.” FairSpeech. Link
Listen to audio examples and compare human speech with machine-generated transcripts.
Reading
Wilson, Kyra, and Aylin Caliskan. “Gender, Race, and Intersectional Bias in AI Resume Screening via Language Model Retrieval.” Brookings Institution, 25 Apr. 2025. Link
Identical resumes with different race- and gender-associated names receive different rankings.
Reading
AlDahoul, Nouar, Talal Rahwan, and Yasir Zaki. “AI-Generated Faces Influence Gender Stereotypes and Racial Homogenization.” Scientific Reports, vol. 15, 2025. Link
Examines racial and gender stereotypes in Stable Diffusion images across 32 professions.
Reading
Omar, Mahmud, et al. “Sociodemographic Biases in Medical Decision Making by Large Language Models.” Nature Medicine, vol. 31, 2025, pp. 1873–1881. Link
Researchers hold clinical cases constant while changing sociodemographic information and compare model recommendations.
Reading
Jonas, Anne, and Jenna Burrell. “Friction, Snake Oil, and Weird Countries: Cybersecurity Systems Could Deepen Global Inequality through Regional Blocking.” Big Data & Society, vol. 6, no. 1, 2019. Link
Examines regional blocking, fraud detection, false positives, and how automated security systems can treat geographically patterned behavior as suspicious.
Reading
Zewe, Adam. “Avoiding Shortcut Solutions in Artificial Intelligence.” MIT News, 2 Nov. 2021. Link
Explains shortcut learning using the example of image classifiers learning to associate cows with grass rather than recognizing the cow itself.
Reading
Geirhos, Robert, et al. “ImageNet-Trained CNNs Are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness.” ICLR, 2019. Link
Shows that image classifiers may rely more strongly on texture than shape; includes striking shape-texture conflict examples.
Topic / Focus
Distinguish training examples from test examples, and separate model performance from the decision to rely on its output. Document one failure and whether revising the training set addresses it.
In This Class
- Is it a Fish Activity
- Key terms:
- False positive
- False negative
- Overfitting
- Training v. Test Data
- How I learned in-class presentations
Slides and Activities
Societal / Ethical Questions
- Whose examples are represented?
- What gets left out?
- Which errors matter most, and to whom?