SYS 478: Fall 2026

Technical Explainer

Supervised Learning

How machines learn from labeled examples – and why what the system learns depends entirely on what people decided to label.

What Is Supervised Learning?

Supervised learning starts with a question like this:

Here are resumes from people who became great employees. Here are resumes from applicants we passed on. If we give you a few thousand resumes of people that we hired / passed on, can you build us a resume screening model – a way of deciding which features of a resume matter – to tell the two groups apart?

This involves having labeled data – which resumes that lead to hirings and which didn’t – and a supervised learning algorithm that can figure out which aspects of a resume matter most.

That is valuable when the same sorting problem repeats at scale – screening medical images, filtering spam, scoring risk. It is limited when labels are noisy, biased, or a poor guide to the future. The hiring example shows why: “great employee” and “passed on” are not facts written in the resume. They are past human decisions. The system learns to reproduce them – including whatever criteria, inconsistencies, or biases those decisions carried.

Example: Medical Diagnostics

One widely studied example of supervised learning is screening for diabetic retinopathy – a disease that can cause vision loss in people with diabetes. It develops slowly, and catching it early matters. But there are not enough eye specialists in many parts of the world to screen every patient who needs it (and of course training a model to do this kind of thing means that fewer people are needed to do this task). Researchers had an idea:

Could a computer learn to look at photographs of the eye and detect early signs of the disease?

To teach it, they collected thousands of eye photographs that specialist doctors had already reviewed and labeled – each image marked with a category: no disease, mild, moderate, severe, or proliferative (the most serious stage).1

The system studied those labeled photographs. It found patterns connecting what the image looks like to what category doctors assigned. Now, when it sees a new photograph it has never encountered before, it can predict which category it belongs to.

How It Works

1. Label thousands of examples

Specialist doctors reviewed thousands of eye photographs and assigned each one a category. This is called the training data – a large collection of examples where the correct answer is already known.

2. Train the system

The system looks at all the labeled photographs and tries to find patterns – combinations of visual features that tend to appear in each category. It adjusts its internal settings based on those patterns until it can predict labels accurately on the examples it has seen.

3. Classify new photographs

When a new, unlabeled photograph comes in, the system assigns it a category – and a confidence score showing how certain it is.

Once trained, the classifier can be used on many new unlabeled eye images, sorting each one into one of the five categories it learned from the labeled examples.

Notably, a very similar technique can be applied to other kinds of classification tasks – whether a system is sorting emails into folders, identifying species of plants from photographs, or estimating how likely a patient is to be readmitted to a hospital. What changes is the number of categories, the meaning of each one, and the stakes involved.

What Labels Are – and Why They Matter

Before the system could learn anything, doctors had to label thousands of photographs. That labeling process looks straightforward – but it involves real judgment calls.

Two specialists looking at the same photograph sometimes disagree. Mild and moderate can be genuinely hard to distinguish. The criteria for each category had to be defined, agreed upon, and applied consistently – by people, with all the variability that involves.2

The system learns to reproduce the labels it was given – not some objective truth about the photographs. If the labeling criteria change, or if different specialists labeled different parts of the dataset, the system inherits that inconsistency.

What Can Go Wrong

Training data carries the context it came from. The diabetic retinopathy system was trained largely on photographs from specific populations and camera equipment. When researchers tested it on photographs from other contexts – different cameras, different populations, different lighting conditions – performance dropped.3 The system had learned patterns from one setting and was being asked to generalize to another.

Labels encode the judgment of the people who assigned them. In medicine, labeling requires expertise and careful protocols. In other contexts – hiring, criminal justice, credit – labels come from historical decisions that may have been discriminatory, inconsistent, or shaped by institutional pressures. The system cannot tell the difference between a carefully considered label and a biased one. It treats them all as ground truth.

A single accuracy number can hide unequal performance. A system that performs well on average may perform worse for patients whose photographs look different from the majority of the training data – different skin tones affecting how blood vessels appear, for example. That variation may not appear in overall accuracy statistics.

The same approach applied elsewhere raises harder questions. Supervised learning trained on historical hiring decisions will learn to reproduce those decisions – including any biases they contained. The same mechanism that helps detect disease can sort job applicants, score loan applications, or predict recidivism. The technique is the same. What changes is what was labeled, by whom, and with what consequences when the system gets it wrong.

Key Takeaways

  1. Supervised learning finds patterns in labeled examples. It does not discover truth on its own.
  2. Labels are human decisions, so the system learns to reproduce judgment, not objective reality.
  3. Training data carries the context it came from, which means a system may work differently in a new setting.
  4. Accuracy alone is not enough. A system can perform well on average while still harming some groups more than others.
  5. The same technique can be used across many domains, but the consequences of being wrong depend on where it is used.

References

Footnotes

  1. Gulshan, V., Peng, L., Coram, M., et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA. 2016;316(22):2402–2410. JAMA full text ↩

  2. Sayres, R., et al. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy. Ophthalmology. 2019;126(9):1264–1272. AAO abstract ↩

  3. Voets, M., Møllersen, K., & Bongo, L. A. Reproduction study using public data of Gulshan et al.’s diabetic retinopathy algorithm. PLOS ONE. 2019;14(6):e0217541. PMC full text; Zhou, Y., et al. Deep learning generalization for diabetic retinopathy staging from fundus images. Physiol Meas. 2025. IOPscience ↩

Back to Technical Explainers