Almost every skin-analysis model on the market today was trained on skin that looks nothing like most of India's. Not through malice — through data availability. The public dermatology datasets the field was built on came overwhelmingly from a handful of Western countries, and the vast majority of those images carry no skin-tone label at all.
The result is a category-wide default: models validated on Fitzpatrick I–II skin first, then extended to darker tones afterwards. For a market where Fitzpatrick III–VI is the norm rather than the exception, that ordering isn't a rounding error. It is the entire distribution being treated as an edge case.
The training-data problem, with numbers
This isn't a vibe. It's one of the most consistently reproduced findings in dermatology AI research. A systematic review of publicly available skin-image datasets published in The Lancet Digital Health found that where a country of origin was recorded at all, images clustered heavily in a small number of regions — roughly 79% traced to just three countries — and only about 1.3% of images carried any ethnicity or skin-tone metadata.
That silence compounds. A model trained on unlabelled data can't be audited by tone. A benchmark built from that data can't report a per-tone breakdown. And a vendor quoting a single headline accuracy number from that benchmark isn't necessarily being dishonest — they may simply have no way to answer the question you actually care about.
What "extended to darker skin tones" actually means in practice
In vendor copy, "now supports all skin tones" usually describes one of three things, and the difference matters:
- Colour normalisation. The pipeline white-balances an image before analysis, so the model sees a more consistent input. Useful — but it corrects for lighting, not for what the model learned.
- Fine-tuning on a small diverse set. A model trained mostly on lighter skin gets a comparatively small top-up of darker-skin images. Measurably better than nothing, and genuinely reported to improve accuracy on Fitzpatrick IV–VI — but the model's underlying priors were set elsewhere.
- Trained and evaluated on III–VI from the start. The test set that decides whether the model ships is majority melanin-rich skin. This is the only version of the claim that changes what the model optimises for.
The first two are the common case. They are worth having. They are also not the same claim as the third, and buyers are rarely told which one they're getting.
What the Fitzpatrick scale is (and isn't)
Built for sunburn risk in 1975, not skin-tone classification
Dermatologist Thomas Fitzpatrick introduced the scale in 1975 to answer a narrow clinical question: how much ultraviolet light can this patient take before they burn? It classifies UV response — burning and tanning behaviour — and it was designed to help dose phototherapy safely. It was never built as a taxonomy of human skin colour.
The industry adopted it anyway, because it was the only widely understood scale available. That's how a phototherapy dosing tool became the default label for training and evaluating computer-vision models.
Why its granularity skews toward lighter tones
Look at how the six categories divide the range and the skew is obvious: four of the six describe skin that burns to some degree, and only two cover the deeply pigmented end. Nearly all of India, and most of South and Southeast Asia, lands in III–VI — compressed into fewer, coarser buckets than lighter skin gets.
| Type | UV response | Typically described as |
|---|---|---|
| I | Always burns, never tans | Very fair |
| II | Burns easily, tans minimally | Fair |
| III | Burns moderately, tans gradually | Light to medium |
| IV | Burns minimally, tans readily | Olive to moderate brown |
| V | Rarely burns, tans profusely | Brown |
| VI | Never burns, deeply pigmented | Deep brown to black |
This is why newer measures exist — the individual typology angle (a computed measure from colourimetry) and the Monk Skin Tone scale, developed specifically to give richer coverage across deeper tones. Neither has displaced Fitzpatrick as the labelling default, which means the field is still evaluating melanin-rich skin with an instrument that was built to do something else.
Why this is India's centre of the distribution, not its edge case
"Edge case" is a statistical description, not a moral one — and it's simply wrong here. Fitzpatrick III–VI isn't a minority segment of the Bharatiya market to be handled after launch. It is the market. A model that performs well on I–II and acceptably on III–VI has, for an Indian D2C brand, been optimised almost entirely for customers that brand does not have.
"Global AI, localised" vs. "built for III–VI first"
The two phrases sound like marketing variants of each other. They describe different engineering decisions:
- Global AI, localised — the model and its evaluation set were fixed elsewhere; localisation adjusts inputs, language, or the product catalogue on top.
- Built for III–VI first — the training data, the labelling panel, and the test set that gates release are all majority melanin-rich skin. Failures on those tones block a release rather than getting logged as a known limitation.
The practical test for any vendor is short: ask for accuracy broken out by Fitzpatrick type. If a per-tone breakdown doesn't exist, the model wasn't evaluated on the question that matters most to a Bharatiya brand.
What this costs a beauty brand, not just an ethics footnote
Skin-tone bias gets written about as a fairness problem. For a D2C beauty brand it also shows up as a straightforwardly commercial one, through a chain that's easy to trace:
- A pigmentation or texture read is under-called on deeper skin, because contrast-based features behave differently against higher melanin.
- The recommendation engine maps that weak read to the wrong products in your catalogue.
- The customer buys, sees no result, and returns — or quietly doesn't repeat.
- Returns rise on exactly the segment that should be converting best, and average order value stays flat despite "personalisation" being live on the site.
The failure is invisible in aggregate metrics. Site-wide conversion looks fine. The read quality on the tones that make up most of your customer base is what moved — and nothing in a standard analytics dashboard surfaces that. If you're weighing this against a quiz-based setup, the mechanism is worth understanding before the spend: see how we think about personalisation for D2C beauty.
What we do differently at Rupam — verified, and still in training
Rupam is built the third way described above: trained and evaluated on Fitzpatrick III–VI Bharatiya skin by default, not retrofitted to it. Concretely, that means the model returns a 14-parameter skin read — acne, pores, dark circles, pigmentation, wrinkles, oiliness, redness, texture, fine lines, eye bags, blackheads, hydration, moles and skin glow — from a single photo, through a REST API and Shopify plugin.
That last line costs us some marketing altitude, and we'd rather pay it. A vendor that won't tell you where its model stops is not giving you enough information to evaluate where it works. Our current numbers, methodology, and open limitations are published in full on the accuracy page.
Frequently asked
- Does the Fitzpatrick scale measure skin tone?
- Not directly. The Fitzpatrick scale was created in 1975 to classify how skin responds to ultraviolet light — how readily it burns or tans — in order to dose phototherapy safely. It became a proxy for skin tone because it was the most widely understood scale available, not because it was designed to measure colour. This is why supplementary measures such as the individual typology angle and the Monk Skin Tone scale now exist.
- Why does skin-AI accuracy vary by skin tone?
- Mainly because of training data. Public dermatology image datasets are drawn overwhelmingly from a small number of Western countries, and only a small fraction of images carry any skin-tone label — so models learn most of their priors from lighter skin and cannot be audited by tone afterwards. Contrast-based visual features also behave differently against higher melanin, so conditions like pigmentation and redness are more easily under-read on deeper skin unless the model was trained and evaluated for it.
- What is Fitzpatrick III–VI skin?
- Fitzpatrick III–VI covers the range from skin that burns moderately and tans gradually (III) through to deeply pigmented skin that never burns (VI). It includes most of the Indian subcontinent as well as much of South and Southeast Asia, Africa, and Latin America. For a Bharatiya D2C beauty brand, Fitzpatrick III–VI is not a minority segment — it is effectively the whole customer base.
Sources
- Systematic review of publicly available skin-image datasets — The Lancet Digital Health, 2022
- Fitzpatrick, T.B. — The validity and practicality of sun-reactive skin types I through VI — Archives of Dermatology, 1988
- Monk Skin Tone scale — a ten-point scale for more representative skin-tone coverage — Google Research, 2022
See how Rupam evaluates against Fitzpatrick III–VI
Our benchmark methodology, current per-tone results, and the limitations we haven't solved yet — published rather than summarised.
View the accuracy report

