Clinical Accuracy of AI Dermatology Tools Versus Clinicians
AI outperforms dermatologists in labs but struggles with real patients and darker skin tones.

What controlled-setting studies actually show about AI versus dermatologist accuracy
If you've followed AI dermatology coverage at all, you've likely seen the headline: AI beats dermatologists at detecting melanoma. In controlled settings, on curated dermoscopic image datasets, this actually holds up. One oft-cited benchmark puts AI at roughly 92.5% accuracy on melanoma detection, against 86.6% for dermatologists working the same images. A 2025 systematic review spanning 18 studies and more than 70,000 test images found pooled AI sensitivity of 0.91 and an AUROC (area under the receiver operating characteristic curve, a measure of overall diagnostic accuracy) of 0.88. Numbers that would satisfy most clinicians.
The specificity figure (the tool's ability to correctly rule out cancer) from that same review is where you have to stop. Pooled specificity came in at 0.64, with a 95% confidence interval stretching from 0.47 to 0.78. That is not a tight result. High sensitivity with low specificity means the tool is good at flagging potential cancers, but it also flags a lot of things that are not cancer: more referrals, more biopsies, more anxiety, more cost. For non-melanoma skin cancers, a 2024 analysis across 44 studies found average AI accuracy around 86.80%, which is respectable but not transformative.
I keep coming back to this: controlled-setting performance is a ceiling, not a floor. And that distinction matters enormously once you start asking what the tool actually does in a hospital corridor at 2pm with a rushed resident.
Where AI shows its clearest advantage: assisting non-specialists at the point of first contact

Here is where the case for AI dermatology tools is most honest and most compelling, because it stops pretending AI replaces clinical judgment and starts asking where it fills a genuine void.
A 2024 meta-analysis of 10 studies tracked what happens to clinician accuracy when AI assistance enters the picture. Without AI, pooled sensitivity among clinicians sat at 74.8%, specificity at 81.5%. With AI: 81.1% and 86.1%. The largest gains landed with non-dermatologists. Separate meta-analytic data shows AI assistance improves primary care provider sensitivity by roughly 13 percentage points, specificity by 11. For a primary care physician in a rural county with no dermatologist within 90 miles, that delta is not a rounding error. That is the difference between catching a melanoma and sending someone home.
The FDA authorized DermaSensor in January 2024, the first AI-enabled device cleared specifically for primary care skin cancer detection. Its pivotal trial across 224 high-risk lesions achieved 95.5% sensitivity. A companion study found that missed skin cancers dropped from 18% to 9% with the tool in use. That is a concrete outcome, not a benchmark score.
The framing that survives scrutiny here is triage and referral, not replacement. AI flags what a non-specialist might miss; a dermatologist makes the final call. The fit between what this technology can actually do and what the clinical moment requires is cleanest in this context, which is why the evidence here is the most persuasive piece of the whole picture.
How AI performance degrades when it leaves the lab
Commercial platforms routinely report accuracy in the low-to-mid 90s in their own validation studies, on curated datasets, under controlled imaging conditions. What happens in a busy NHS trust is considerably more sobering, and this gap gets papered over more than it should be.
A real-world NHS evaluation of one triaging tool found specificities for detecting cancer lesions ranging only from 70.1% to 73.4% across two trusts. A second-read reviewer overturned 40 to 50% of the cases the tool had flagged for discharge. That is not a minor calibration issue. That is a structural gap between what a tool does in validation and what it does with real patients.
Why does it degrade? Because AI sees an image. A dermatologist integrates patient history, lesion evolution over time, physical examination findings, and the kind of gestalt pattern recognition that develops over years of seeing thousands of unusual presentations. Atypical melanomas, amelanotic lesions (melanomas that lack the usual dark pigmentation), patients with multiple dysplastic nevi (atypical moles): these are the hard cases, and they are hard precisely because the diagnostic signal is distributed across sources no camera captures. Existing systems achieve only around 68% accuracy on early-stage nodular melanoma, which is the presentation where catching it early matters most.
A 2025 WHO digital health assessment identified persistent structural gaps in commercial platforms: fragmentation between visual analysis and clinical reporting workflows, inadequate patient education components, and binary confidence metrics that strip away clinical nuance. One Medscape analysis in 2024 put it plainly: the translation and integration of AI into actual clinical workflows to benefit patients beyond academic research has been limited.
The lab-to-clinic gap is not a footnote. It is the central unresolved problem in this field, and acknowledging it is the only honest starting point for anyone evaluating these tools.
The skin-tone bias problem and what it means for who benefits

A 2023 systematic review of 232 studies found AI achieving around 90% overall accuracy on cutaneous malignancy detection. That sounds reassuring until you read the methodological fine print: only 1.3% of those studies described Fitzpatrick skin type (a scale from I, very fair, to VI, deeply pigmented), and only 3.2% of training images represented Fitzpatrick types IV through VI.
Research on GPT-4o showed significantly lower sensitivity, specificity, and accuracy for melanoma detection in darker skin tones. A 2024 Nature Medicine study found that AI decision support improved overall accuracy across both dermatologists and primary care providers, but for PCPs specifically, AI assistance actually exacerbated accuracy disparities for darker skin tones by 5 percentage points, a result that reached statistical significance.
The ISIC benchmark dataset, the most widely used source for training dermatology AI, skews heavily toward fair skin, often exceeding 70% representation of lighter tones. A 2025 cross-sectional study examining 4,000 AI-generated dermatologic images found that only about 10% reflected dark skin, and only 15% of those images accurately depicted the intended condition.
This is the structural problem the data keeps circling back to: the populations with the least specialist access are also the populations for whom AI currently performs least reliably. The access argument, the argument that animated the DermaSensor approval and a lot of the field's aspirational energy, does not hold uniformly across the patients who need it most. When researchers trained AI models on more racially and phenotypically diverse datasets, accuracy for Fitzpatrick types IV through VI improved. The problem is solvable. It just has not been systematically solved, and the distance between what is solvable and what has actually been solved is where patients get hurt.
What happens when dermatologists and AI work together

The consistent finding across systematic reviews from 2024 and 2025 is that dermatologists working with AI outperform either operating alone. AI surfaces missed early melanomas; clinicians correct AI errors. That is the shape of the collaboration that the data supports.
But there is an asymmetry embedded in that finding that is worth sitting with. Experienced dermatologists are less influenced by incorrect AI suggestions; they have enough accumulated pattern recognition to notice when the model is wrong. Non-dermatologists, who represent precisely the setting where the access gap is most acute, are more likely to follow incorrect AI outputs. When the AI makes a mistake and a clinician follows its lead, the clinician also makes that mistake. Error alignment is documented in human-AI interaction research, not theoretical.
This reframes the design question considerably. The issue is not simply whether AI improves mean accuracy in aggregate. It is whether the interface and workflow are structured so that the AI's strengths cover clinician blind spots without propagating its own failure modes through the exact populations who have no specialist backstop. That is an interface design problem, a workflow problem, and a training problem as much as it is an algorithmic one.
The liability and transparency questions that remain unresolved at the clinical level
When clinicians are surveyed about AI in dermatology, three concerns surface consistently: divestment of healthcare decision-making to large technology companies, medical liability, and diminishing reliance on specialists. Liability for machine error ranks first among dermatologists. That is not a trivial preoccupation with legal exposure; it reflects a genuine structural problem.
As of 2025, the AAD and major international dermatology organizations had not issued position statements specifying how liability should be allocated when an AI diagnostic error causes patient harm. The FDA had approved three AI-powered dermatology spectroscopy devices, but no FDA-approved AI-based dermatology mobile application exists. Most direct-to-consumer AI skin apps operate without regulatory clearance or sufficient clinical evidence supporting their accuracy claims.
The black box problem compounds all of this. A dermatologist who makes a diagnostic error can articulate the cognitive chain that led there; an audit can reconstruct where the reasoning broke down. AI systems typically cannot offer that. The opacity makes error attribution difficult, which matters enormously when a missed melanoma ends up in litigation and no one can explain what the algorithm weighted or why.
DermaSensor's FDA authorization surfaces a question the field has not answered: what ongoing evidence generation should be required after a device clears regulatory review? Initial approval reflects pivotal trial performance; real-world performance, as the NHS evaluation demonstrates, can look quite different. Accuracy benchmarks, however strong, do not resolve interpretability, liability allocation, or workflow integration. Those are the friction points that determine whether a technically accurate tool ever actually reaches the patient.
What the evidence adds up to for patients and the clinicians evaluating these tools
I want to be direct about where I come out on this, because I think the field has a habit of resolving these questions too cleanly in both directions.
In controlled settings on standard dermoscopic images, AI accuracy is competitive with dermatologists. That part of the enthusiasm is not hype. The strongest real-world evidence lives in primary care triage, AI as referral decision-support for non-specialists, where the clinical gap is largest and the outcome data, halving missed cancer rates in the DermaSensor trial, is most credible.
Real-world deployment consistently underperforms controlled benchmarks. That gap has not been systematically closed, and the clinicians and patients on the wrong side of it are not well served by pretending otherwise.
Skin-tone bias is a structural limitation in the current evidence base, not a peripheral concern. Tools validated predominantly on lighter skin tones cannot be assumed to generalize, which means the populations most underserved by specialist access are also the populations for whom the AI promise is most unreliable. That is an uncomfortable finding for a field that leans heavily on the equity argument.
The collaboration model is the most evidence-supported frame, but it requires deliberate interface design to avoid propagating model errors, particularly when non-specialists are operating without a specialist safety net nearby.
None of this points toward replacement. What the evidence points toward is structured support: AI-generated assessment integrated into a care pathway that keeps a clinician in the loop, with explicit attention to whose skin is in the training data, how errors surface and get audited, and who bears responsibility when the system fails. Nolla, an AI-assisted skincare telehealth service that routes patients to licensed US clinicians for diagnosis and prescriptions in roughly 10 minutes, is one example of that model applied to everyday dermatology questions. The field is not there yet. But that is the gap worth naming, because naming it precisely is how it eventually gets closed.
Sources
- AI vs Dermatologists: Who Detects Melanoma Better in 2025? Latest Research & Results
- The Use of Artificial Intelligence for Skin Disease Diagnosis in Primary Care Settings: A Systematic Review - PMC
- Diagnostic accuracy of artificial intelligence compared to family physicians and dermatologists for skin conditions: a systematic review and meta-analysis - PMC
- Evaluation of the Accuracy of Artificial Intelligence (AI) Models in Dermatological Diagnosis and Comparison With Dermatology Specialists - PMC
- Equity and Generalizability of Artificial Intelligence for Skin-Lesion Diagnosis Using Clinical, Dermoscopic, and Smartphone Images: A Systematic Review and Meta-Analysis
- Artificial Intelligence in Skin Cancer Diagnosis: A Reality Check - ScienceDirect


