"AI is biased against darker skin" gets said often enough that it risks sounding like a slogan rather than a finding. It shouldn't — the underlying research is specific, peer-reviewed, and has been reproduced across multiple independent studies over the past several years.
This is a synthesis of what that research actually shows: where the training-data imbalance comes from, why the field's own labelling tool has built-in limits, what the accuracy gap looks like in numbers, and what measurably closes it. Every claim below is attributed to a named source.
The training-data imbalance
The starting point for most dermatology AI models is a public image dataset, and those datasets are not representative of the world's skin tones. A systematic review of 70 publicly available skin-image datasets, published in The Lancet Digital Health, found that country of origin was recorded for only a fraction of images — and where it was, the images clustered heavily in a small number of Western countries. Skin-tone or ethnicity metadata was present for roughly 1.3% of images across the datasets reviewed.
That matters beyond a fairness argument: a model can only be evaluated on what its data recorded. A dataset that never labelled skin tone cannot produce a benchmark that reports accuracy by skin tone — the gap in the label becomes a gap in what can even be audited later.
The Fitzpatrick scale's own limitations as a labelling tool
Where datasets do carry a skin-tone label, that label is almost always the Fitzpatrick scale — a system dermatologist Thomas Fitzpatrick introduced in 1975 to classify UV burning and tanning response for phototherapy dosing, not skin colour. Research published in npj Digital Medicine examining its use as an AI labelling standard notes that its six categories were never validated as a colour taxonomy, and that self-reported Fitzpatrick type in particular correlates poorly with measured skin reflectance.
A 2026 classifier study in the Journal of the American Academy of Dermatology found similar instability: models trained to predict Fitzpatrick type from images showed meaningfully lower agreement with dermatologist ratings at the darker end of the scale (types V–VI) than at the lighter end — the label itself is least reliable exactly where accurate labelling matters most for Fitzpatrick III–VI populations.
What the accuracy-gap research actually shows
The downstream effect shows up as a measurable sensitivity gap in clinical-task benchmarks, most consistently in melanoma and pigmented-lesion detection. Independent evaluations comparing model performance across Fitzpatrick groupings have repeatedly found lower sensitivity and specificity on Fitzpatrick IV–VI images than on Fitzpatrick I–II images evaluated by the same model — the direction of the gap is consistent across studies even where the exact magnitude varies by dataset and task.
This isn't unique to melanoma classifiers. The same pattern — degraded performance correlated with higher Fitzpatrick type or darker measured skin tone — has been documented across broader dermatology image-classification tasks, which is consistent with a training-data explanation rather than a disease-specific one.
What measurably improves it
The research is not only diagnostic — it also points at what closes the gap, and it is less exotic than it might sound. Studies that retrain or fine-tune models on deliberately tone-diverse datasets report improved sensitivity on Fitzpatrick IV–VI images, in some cases closing a meaningful share of the gap measured against the same model's baseline performance. The improvement tracks with how representative the added data actually is, not just its volume — a small amount of well-labelled, expert-reviewed diverse data has moved accuracy more than a large amount of unlabelled or self-reported data in several of these evaluations.
- Representative training data, not just more data. Volume alone doesn't close the gap — several studies show a smaller, well-labelled diverse dataset outperforming a larger unlabelled one on darker-tone accuracy.
- Expert-reviewed labels over self-reported Fitzpatrick type. Self-reported skin type introduces the same instability documented in the JAAD classifier study above; dermatologist-graded or colorimeter-measured labels are more reliable.
- Evaluation sets that gate release by tone, so a model can't ship with an undetected gap on Fitzpatrick IV–VI simply because aggregate accuracy looked acceptable.
What this means for a product decision, not just a research summary
For a brand or buyer evaluating a skin-AI vendor, the research above converts into a short, concrete question: was this model trained and evaluated on a dataset that actually labelled and represented Fitzpatrick III–VI skin, or was that label largely absent from the data it learned from? A vendor that can answer with a per-tone accuracy breakdown is describing a model built the way the research says closes the gap. A vendor that can't is, per the studies above, describing a model that was most likely never tested against the question.
This is the same underlying mechanism covered from the India-market angle in our piece on why Fitzpatrick III–VI is still treated as an edge case in most skin AI — that article applies the research below to what it costs a D2C beauty brand specifically; this one stays with the research itself.
Sources
Full citations for every claim above, for anyone who wants to verify or build on this directly:
- Systematic review of 70 publicly available skin-image datasets — The Lancet Digital Health, 2022.
- Examination of the Fitzpatrick scale as an AI labelling standard, including self-report vs. measured reflectance — npj Digital Medicine.
- Fitzpatrick-type classifier agreement study, showing lower dermatologist-agreement at darker skin types — Journal of the American Academy of Dermatology, 2026.
- Fitzpatrick, T.B. — The validity and practicality of sun-reactive skin types I through VI, Archives of Dermatology, 1988.
Frequently asked
- Is AI bias in dermatology a real, documented problem?
- Yes. Multiple independent, peer-reviewed studies — including a systematic review of 70 public skin-image datasets in The Lancet Digital Health and classifier-agreement research in the Journal of the American Academy of Dermatology — document both a severe skin-tone labelling gap in training data and a consistent accuracy gap between lighter and darker Fitzpatrick types in resulting models.
- Why do dermatology AI models perform worse on darker skin tones?
- Primarily because the public datasets most models are trained on carry skin-tone labels for only a small fraction of images and cluster heavily toward a handful of Western countries, so the model's priors are set mostly by lighter-skin imagery. Compounding this, the Fitzpatrick scale most datasets use for labelling was built to classify UV burn response, not skin colour, and is documented to be least reliable at the darker end of the scale — exactly where accurate labelling matters most.
- What actually improves skin-AI accuracy on darker skin?
- Research points to deliberately tone-representative training data, evaluated with expert-reviewed rather than self-reported skin-tone labels, and release gated on per-tone accuracy rather than an aggregate score. Studies retraining models on diverse, well-labelled data report meaningful sensitivity improvements on Fitzpatrick IV–VI images — and the gain tracks with how representative the added data is, not simply how much of it there is.
Sources
- Systematic review of publicly available skin-image datasets — The Lancet Digital Health, 2022
- The Fitzpatrick scale as an AI labelling standard — self-report vs. measured reflectance — npj Digital Medicine, 2023
- Fitzpatrick-type classifier agreement across skin types — Journal of the American Academy of Dermatology, 2026
- Fitzpatrick, T.B. — The validity and practicality of sun-reactive skin types I through VI — Archives of Dermatology, 1988
See how we evaluate for this specific gap
Our per-tone benchmark methodology, current results, and the limitations we haven't solved yet — published rather than summarised.
View the accuracy report

