Skin Comparisons
Skin, examined

FDA Regulation and Clinical Validation of AI Skin Tools

Most AI skin tools cleared by the FDA took the easiest regulatory route, not the most rigorous one.

Staff Writer · · 12 min read
Cover illustration for “FDA Regulation and Clinical Validation of AI Skin Tools”
AI-Powered Skin Analysis · September 9, 2026 · 12 min read · 2,589 words

"FDA cleared" gets treated like a stamp of safety, a green light that means a device works and works well. But that phrase describes a process, not a performance guarantee. Knowing which process a device went through, and what that process actually demanded of the company that built it, changes how much weight the phrase should carry.

Start with the legal trigger. Under Section 201(h) of the Federal Food, Drug, and Cosmetic Act, FDA treats something as a medical device if its intended use is diagnosis, cure, mitigation, treatment, or prevention of disease. That's a broad net. An app that looks at a photo of a mole and says "this looks concerning" falls under that definition just as much as a piece of imaging hardware in a hospital.

Software as a Medical Device, or SaMD, covers standalone software that runs on regular platforms like phones, laptops, or cloud servers, and gets regulated under the same framework as physical devices. Almost every AI skin tool on the market today counts as SaMD.

There are three main doors a device can walk through to get cleared. 510(k) clearance asks whether a new device is substantially equivalent to something already on the market, a predicate device. De Novo classification applies to novel devices that are low-to-moderate risk but have no existing predicate to compare against. Premarket Approval, or PMA, is the strictest path, reserved for high-risk devices, and it demands the most clinical evidence.

That distinction matters more than it sounds like it should. Clearance under 510(k) means the FDA reviewed a submission and judged it similar enough to something already approved. It does not mean an independent study proved the device beats the standard of care, or even that a randomized trial happened at all. Keep that gap in mind. The next section shows just how much of the market runs through that lighter door.

How 97% of cleared AI devices traveled the 510(k) path, and what that pathway does and doesn't require

The number of AI-enabled devices with FDA authorization has climbed fast: over 1,250 as of July 2025, up from 950 in August 2024, and fewer than 400 back in 2020. That's roughly triple in under five years.

Of the devices cleared as of August 2024, 97% went through 510(k). To repeat what that pathway asks: not "does this work better than what doctors do now," but "is this substantially equivalent to a device already cleared." A company can win clearance by showing their AI tool behaves similarly to an older, already-cleared tool. Proving it improves outcomes isn't part of the deal.

The rest of the landscape is small by comparison. Twenty-two devices went through De Novo, the path for genuinely new devices with no predicate to point to. Only 4 devices needed the full rigor of PMA, the track built for high-risk devices where the stakes of getting it wrong are highest.

So when a device says "FDA cleared," ask which of these three doors it walked through. A tool cleared via 510(k) may have shown it resembles an existing device, not that it outperforms a dermatologist's exam or even a simpler diagnostic method already in clinics.

Research examining cleared AI tools has found persistent gaps: unclear evidence of real clinical benefit, unclear generalizability, and public FDA filings that often left out basic details like how testing was done, what validation cohorts looked like, or what steps, if any, were taken to check for bias. None of this means 510(k) clearance is broken or that cleared devices are dangerous. It means the label alone doesn't tell a patient or clinician what was actually tested, and what wasn't.

Diagram: Three Doors to FDA Clearance — and How Few Devices Use the Hardest One. Visualizes: Show the three FDA clearance pathways as a ranked or funnel-style breakdown by device count, emphasizing the extreme concentration at the lightest end.

The grey zone: consumer skin apps that claim to diagnose without any FDA authorization

Here's the line that matters: if an app tells someone their mole has a certain probability of being cancerous, or says they "may have melanoma," that's a medical device claim. It doesn't matter what the company calls the product in its marketing copy or App Store listing. The FDA has already sent warning letters to companies operating exactly in this space.

The scale of it is bigger than most people assume. An analysis of 95,731 apps identified as likely SaMD on the Google Play Store found that 95% had no verifiable evidence of CE marking in the EU or FDA authorization in the US. That's not a small loophole. That's most of the market operating with zero verified regulatory review behind it.

What makes this dangerous isn't obvious carelessness. It's the opposite: a clean interface, confident-sounding output, maybe a percentage score that looks clinical, and then, buried somewhere in the terms, a line saying "not intended to diagnose." The polish does the convincing. The disclaimer does the legal work.

Radiology, for context, makes up 76% of all FDA-listed AI devices. Dermatology is a smaller, faster-growing slice of that pie, which creates an odd mismatch: this is a field where consumer apps are multiplying quickly, but where actual cleared devices remain few. Knowing what a cleared device had to go through gives a patient or clinician a real yardstick for telling the authorized apart from the merely confident.

DermaSensor as a case study in what FDA authorization of an AI skin device actually required

On January 17, 2024, the FDA authorized DermaSensor, the first AI-enabled device cleared for skin cancer detection in a primary care setting. It had already been authorized in the EU, Australia, and New Zealand before reaching the US market.

DermaSensor is cleared to help detect all three common skin cancers: melanoma, basal cell carcinoma, and squamous cell carcinoma. The mechanism is worth pausing on, because it's not what most people picture when they hear "AI skin tool." It doesn't analyze a photo. It uses spectroscopy, pulsing light into a lesion and having the AI assess cellular characteristics from the reflected signal. It was built for primary care doctors, not dermatologists, meaning the whole design premise is putting a screening tool in the hands of non-specialists.

The pivotal trial behind it, DERM-SUCCESS, ran across 22 centers and enrolled more than 1,000 patients. The device hit 96% sensitivity across all three cancer types. When primary care physicians used it as a decision-support tool rather than relying on their own judgment alone, the rate of missed skin cancers dropped by roughly half.

That's a meaningful regulatory first: an AI dermatology device cleared specifically for use by non-specialists. But analysis in npj Digital Medicine noted that how the device performs across the full range of real-world clinics, patient populations, and lighting conditions is a separate question that ongoing evidence collection needs to answer.

Zoom out and the field is still narrow. A 2025 review in the International Journal of Dermatology counted 15 regulatory-approved AI dermatology devices worldwide, with only 3 FDA-approved systems in the US, focused on melanoma and skin cancer detection. One more distinction worth flagging: DermaSensor is SiMD, Software in a Medical Device, meaning the software is tied to a specific physical instrument. That's different from pure SaMD, standalone software running on a phone or in the cloud. Anyone comparing DermaSensor to an app-based tool is comparing two different categories of device, not two versions of the same thing.

What clinical validation studies actually measure, and where the performance numbers come from

A systematic review and meta-analysis that searched literature through October 2025, pulling from tens of thousands of test images, gives a useful baseline: pooled sensitivity of 0.91, specificity of 0.64, and an AUROC of 0.88.

That specificity number deserves a second look. A pooled 0.64 with a confidence interval spanning 0.47 to 0.78 means results varied widely depending on the study, the setting, and the type of lesion being tested. High sensitivity sounds reassuring, since it means the tool catches most true cancers. But without specificity to match, a lot of what it catches includes things that aren't cancer at all: more false positives, more unnecessary biopsies, more anxiety for patients who didn't need any of it.

How do these tools stack up against actual dermatologists? Recent reviews have found AI models reaching diagnostic accuracy comparable to dermatologists in controlled settings. But most of the underlying comparison studies looked at a narrow range of conditions, and the gap between controlled study inputs and the full clinical picture a real doctor would draw on remains a recurring limitation.

There's a deeper issue buried in how these studies get built. Most of the reviewed research tested its final model on internal datasets, meaning data pulled from the same institution that trained the model in the first place. A model that shines on its own training population isn't guaranteed to hold up once it meets patients, lighting, and cameras it's never seen before.

A smaller, more recent example makes the scale problem concrete. DermFlow, a dermatology-trained multimodal language model validated at Indiana University between February 2023 and May 2025, was tested on 59 patients and 68 biopsy-proven lesions. It hit 92.6% accuracy on any diagnosis, with 93.9% sensitivity and 89.5% specificity. Strong numbers, but a sample that small, in one setting, doesn't tell anyone how the model behaves at scale.

Underneath all of this sits a basic mismatch: the crisp, curated clinical images used in validation studies aren't what a smartphone camera captures in a dim bathroom with a shaky hand. So when a company advertises an accuracy figure, worth asking: tested on whose data, in what setting, against what benchmark, and across which skin tones?

How AI skin tools perform worse on darker skin tones, and why current validation standards don't always catch it

That last question, skin tone, turns out to be where some of the sharpest gaps show up. A 7-point gap in a diagnostic tool is not a rounding error.

And the problem isn't confined to diagnosis. It shows up even in how AI generates images of skin conditions. A 2025 study evaluating 4,000 AI-generated dermatological images found only 10.2% reflected dark skin, and just 15% accurately depicted the condition they were supposed to show. Broken down by tool, images with Fitzpatrick scores above IV showed up in only 6.0% of ChatGPT-4o outputs, 3.9% of Midjourney's, and 8.7% of Stable Diffusion's.

Why does this keep happening? Trace it back to the training data. These models learn overwhelmingly from images of lighter skin, so their internal sense of what a "typical" lesion looks like skews that direction, too. The literature backs this up from the other side as well: when models are trained on more diverse datasets, their accuracy on Fitzpatrick IV through VI cases improves.

Here's where the regulatory story connects back to the first section. 510(k) clearance does not require a company to publish, or even submit, a breakdown of how a device performs across different skin tones. A device can clear the FDA bar without ever demonstrating it works equally well on Fitzpatrick VI skin as it does on Fitzpatrick II skin.

That's not an abstract fairness concern. As these tools become more available, particularly in settings meant to expand access, gaps like this risk landing hardest on the same patients who already face the biggest barriers to dermatology care. One useful signal to watch for: does the company publish subgroup performance by Fitzpatrick type? If that data is nowhere to be found, treat the absence itself as information.

How the FDA's evolving guidance attempts to address what initial clearance doesn't cover

Traditional FDA review was designed around hardware: a device gets built, tested, cleared, and then stays exactly as it was cleared. AI doesn't sit still that way. A model can keep learning, keep shifting, and drift from the version that earned clearance in the first place. The FDA has spent the past few years building guidance to catch up with that reality.

The timeline reads like a slow tightening of the framework. Good Machine Learning Practice guiding principles arrived in October 2021. Guiding principles for Predetermined Change Control Plans followed in October 2023. Transparency guidance for machine-learning-enabled devices came in June 2024, alongside draft guidance on change control plans in August 2024. December 2024 brought recommendations specifically for AI device software submissions, and January 2025 rounded it out with lifecycle management guidance for AI-enabled device functions.

The centerpiece of this is the Predetermined Change Control Plan, or PCCP, which lets a company pre-specify how its AI model is allowed to change after clearance, within boundaries the FDA signs off on ahead of time. Roughly 10% of 2025 clearances included a PCCP. That's real movement, but it's still a minority. Most cleared devices today have no pre-approved plan for how, or whether, their model can evolve after hitting the market.

Structurally, the FDA has been building out how it coordinates AI oversight across the agency. It published a cross-center approach to AI in March 2024, then in 2025 stood up two new bodies, an External Policy Council and an Internal Use Council. The Digital Health Advisory Committee held its first meeting in November 2024, with a second scheduled for November 2025. All of this is new enough that its effects are still unfolding.

Researchers examining this space have made a pointed argument: real-time monitoring, transparency, and bias mitigation all still have meaningful weaknesses, with calls for mandated post-market evaluation and enforceable fairness standards, not just voluntary guidance. Worth remembering: a device cleared in 2022 wasn't held to standards the FDA only wrote in 2024 or 2025. The date on a device's clearance says something about which rulebook it was actually judged against.

What clinicians and patients can actually do with this information when evaluating an AI skin tool

None of this is a reason to distrust AI skin tools outright. It's a reason to know what questions to ask before trusting one. Start with the pathway: did the device go through 510(k), De Novo, or PMA? That single fact says a lot about how much scrutiny it received.

From there, dig a bit further. Was the clinical validation tested on an independent dataset, or just on data from the same institution that built the model? Does the evidence package break down performance by Fitzpatrick skin type, or does it stay silent on that question? Is there a Predetermined Change Control Plan in place, meaning the FDA has pre-approved how the model is allowed to change over time? And was validation done under conditions that resemble how the tool will actually get used, meaning ordinary image quality, a real mix of patients, not a lab-perfect setup?

For consumer apps with no cleared claim behind them: if a tool assigns a cancer probability without FDA authorization, what it produces isn't a medical assessment. It might be interesting. It isn't a diagnosis.

The role AI should play here is support, not replacement. A good tool nudges a patient or physician toward the right next step, a referral, a biopsy, a follow-up visit, rather than trying to stand in for a clinical exam. A tool that delivers a diagnostic probability score directly to a consumer operates differently from one integrated into a supervised clinical workflow with no one checking the work.

Understanding how a tool was validated, and what it was and wasn't tested on, isn't a technical skill reserved for engineers. It's the same kind of literacy that helps a patient read a treatment plan or push back on a diagnosis that doesn't feel right. Skin health is medical health, plain and simple, and the rigor patients expect from any other diagnostic tool shouldn't lower just because this one runs on a phone.

Sources

  1. FDA Oversight: Understanding the Regulation of Health AI Tools • Bipartisan Policy Center
  2. The illusion of safety: A report to the FDA on AI healthcare product approvals
  3. Learnings from the first AI-enabled skin cancer device for primary care authorized by FDA | npj Digital Medicine
  4. AI Skin Diagnostics in 2026: What's Real, What's Hype, and What Your Device Actually Needs | Social Life Magazine
  5. Artificial Intelligence in Dermatology: A Comprehensive Review of Approved Applications, Clinical Implementation, and Future Directions - Nahm - 2025 - International Journal of Dermatology - Wiley Online Library
  6. Understanding FDA regulations for AI in SaMD | ICON
  7. onlinelibrary.wiley.com
  8. FDA Pathways for AI SaMD: 510(k), De Novo & PMA Guide | IntuitionLabs

More in AI-Powered Skin Analysis