SYS 478: Fall 2026

Thu, Sep 17

Topic 2. How Computers Work

Learning From Data: Training Sets, Errors, and Bias

Please complete the tasks and readings listed below before class on Thu, Sep 17.

Required Readings

Reading
Intro to Supervised Learning (Course Website) Link
Reading
Zewe, Adam. “Can Machine-Learning Models Overcome Biased Datasets?” MIT CSAIL, 2 Mar. 2022. Link

Optional Readings

Reading
Google. Teachable Machine [Interactive]. Link
Try object classification; test unfamiliar backgrounds and examples.
Reading
U.S. Food and Drug Administration. “FDA Proposes Updated Recommendations to Help Improve Performance of Pulse Oximeters Across Skin Tones.” FDA, 6 Jan. 2025. Link
Measurement bias; consider how errors can enter a system before machine learning begins.
Reading
Koenecke, Allison, et al. “Automated Speech Recognition Less Accurate for Blacks.” Stanford Report, 23 Mar. 2020. Link
Interactive examples show substantially higher speech-recognition error rates for Black speakers.
Reading
Stanford Computational Policy Lab. “The Race Gap in Speech Recognition Technology.” FairSpeech. Link
Listen to audio examples and compare human speech with machine-generated transcripts.
Reading
Wilson, Kyra, and Aylin Caliskan. “Gender, Race, and Intersectional Bias in AI Resume Screening via Language Model Retrieval.” Brookings Institution, 25 Apr. 2025. Link
Identical resumes with different race- and gender-associated names receive different rankings.
Reading
AlDahoul, Nouar, Talal Rahwan, and Yasir Zaki. “AI-Generated Faces Influence Gender Stereotypes and Racial Homogenization.” Scientific Reports, vol. 15, 2025. Link
Examines racial and gender stereotypes in Stable Diffusion images across 32 professions.
Reading
Omar, Mahmud, et al. “Sociodemographic Biases in Medical Decision Making by Large Language Models.” Nature Medicine, vol. 31, 2025, pp. 1873–1881. Link
Researchers hold clinical cases constant while changing sociodemographic information and compare model recommendations.
Reading
Jonas, Anne, and Jenna Burrell. “Friction, Snake Oil, and Weird Countries: Cybersecurity Systems Could Deepen Global Inequality through Regional Blocking.” Big Data & Society, vol. 6, no. 1, 2019. Link
Examines regional blocking, fraud detection, false positives, and how automated security systems can treat geographically patterned behavior as suspicious.
Reading
Zewe, Adam. “Avoiding Shortcut Solutions in Artificial Intelligence.” MIT News, 2 Nov. 2021. Link
Explains shortcut learning using the example of image classifiers learning to associate cows with grass rather than recognizing the cow itself.
Reading
Geirhos, Robert, et al. “ImageNet-Trained CNNs Are Biased Towards Texture; Increasing Shape Bias Improves Accuracy and Robustness.” ICLR, 2019. Link
Shows that image classifiers may rely more strongly on texture than shape; includes striking shape-texture conflict examples.