1. Label thousands of examples

Specialist doctors reviewed thousands of eye photographs and assigned each one a category. This is called the training data – a large collection of examples where the correct answer is already known.
Technical Explainer
How machines learn from labeled examples – and why what the system learns depends entirely on what people decided to label.
Supervised learning starts with a question like this:
Here are resumes from people who became great employees. Here are resumes from applicants we passed on. If we give you a few thousand resumes of people that we hired / passed on, can you build us a resume screening model – a way of deciding which features of a resume matter – to tell the two groups apart?
This involves having labeled data – which resumes that lead to hirings and which didn’t – and a supervised learning algorithm that can figure out which aspects of a resume matter most.
That is valuable when the same sorting problem repeats at scale – screening medical images, filtering spam, scoring risk. It is limited when labels are noisy, biased, or a poor guide to the future. The hiring example shows why: “great employee” and “passed on” are not facts written in the resume. They are past human decisions. The system learns to reproduce them – including whatever criteria, inconsistencies, or biases those decisions carried.
One widely studied example of supervised learning is screening for diabetic retinopathy – a disease that can cause vision loss in people with diabetes. It develops slowly, and catching it early matters. But there are not enough eye specialists in many parts of the world to screen every patient who needs it (and of course training a model to do this kind of thing means that fewer people are needed to do this task). Researchers had an idea:
Could a computer learn to look at photographs of the eye and detect early signs of the disease?
To teach it, they collected thousands of eye photographs that specialist doctors had already reviewed and labeled – each image marked with a category: no disease, mild, moderate, severe, or proliferative (the most serious stage).1
The system studied those labeled photographs. It found patterns connecting what the image looks like to what category doctors assigned. Now, when it sees a new photograph it has never encountered before, it can predict which category it belongs to.


Once trained, the classifier can be used on many new unlabeled eye images, sorting each one into one of the five categories it learned from the labeled examples.
Notably, a very similar technique can be applied to other kinds of classification tasks – whether a system is sorting emails into folders, identifying species of plants from photographs, or estimating how likely a patient is to be readmitted to a hospital. What changes is the number of categories, the meaning of each one, and the stakes involved.
Before the system could learn anything, doctors had to label thousands of photographs. That labeling process looks straightforward – but it involves real judgment calls.
Two specialists looking at the same photograph sometimes disagree. Mild and moderate can be genuinely hard to distinguish. The criteria for each category had to be defined, agreed upon, and applied consistently – by people, with all the variability that involves.2
The system learns to reproduce the labels it was given – not some objective truth about the photographs. If the labeling criteria change, or if different specialists labeled different parts of the dataset, the system inherits that inconsistency.
Training data carries the context it came from. The diabetic retinopathy system was trained largely on photographs from specific populations and camera equipment. When researchers tested it on photographs from other contexts – different cameras, different populations, different lighting conditions – performance dropped.3 The system had learned patterns from one setting and was being asked to generalize to another.
Labels encode the judgment of the people who assigned them. In medicine, labeling requires expertise and careful protocols. In other contexts – hiring, criminal justice, credit – labels come from historical decisions that may have been discriminatory, inconsistent, or shaped by institutional pressures. The system cannot tell the difference between a carefully considered label and a biased one. It treats them all as ground truth.
A single accuracy number can hide unequal performance. A system that performs well on average may perform worse for patients whose photographs look different from the majority of the training data – different skin tones affecting how blood vessels appear, for example. That variation may not appear in overall accuracy statistics.
The same approach applied elsewhere raises harder questions. Supervised learning trained on historical hiring decisions will learn to reproduce those decisions – including any biases they contained. The same mechanism that helps detect disease can sort job applicants, score loan applications, or predict recidivism. The technique is the same. What changes is what was labeled, by whom, and with what consequences when the system gets it wrong.
Gulshan, V., Peng, L., Coram, M., et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA. 2016;316(22):2402–2410. JAMA full text ↩
Sayres, R., et al. Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy. Ophthalmology. 2019;126(9):1264–1272. AAO abstract ↩
Voets, M., Møllersen, K., & Bongo, L. A. Reproduction study using public data of Gulshan et al.’s diabetic retinopathy algorithm. PLOS ONE. 2019;14(6):e0217541. PMC full text; Zhou, Y., et al. Deep learning generalization for diabetic retinopathy staging from fundus images. Physiol Meas. 2025. IOPscience ↩