AI matches actives to skin by converting a photo and questionnaire into a numeric profile, then scoring candidate ingredients against it under chemical constraints. It narrows guesswork; it does not remove testing. Published evidence is thin: most quoted accuracy figures come from vendors, and dermatology models lose measurable accuracy on darker skin tones.
- 1,156,703 women analysed by an AI algorithm on smartphone selfies plus questionnaire — acne severity fell after age 25, reached a minimum at 40–44 and rose again (peer-reviewed study, Skin Research and Technology, 2024). Population-scale pattern finding is what this technology demonstrably does well.
- 5 % niacinamide worked, 2 % did not for facial hyperpigmentation at 4 and 8 weeks in split-face trials collected in a 2021 review — yet 2 % niacinamide significantly lowered sebum excretion in Draelos et al. (2006). The effective concentration depends on the endpoint, not on the ingredient alone. Percentages here are active matter unless stated otherwise.
- Dermatology AI models degrade on dark skin tones — demonstrated on the Diverse Dermatology Images (DDI) set, the first pathology-confirmed public dataset with diverse skin tones; fine-tuning on those images closed the gap (peer-reviewed study, Science Advances, 2022).
- 2,737 metagenome-assembled genomes across a multi-centre facial survey produced just two reproducible "cutotypes" — data-driven skin groupings that do not map onto dry/oily/combination at all (peer-reviewed study, Microbiome, 2024).
- No concentrations appear on any label. Under Article 19 of Regulation 1223/2009 as retained in UK law, ingredients below 1 % may be listed in any order and no percentage is declared — so no model can read a formula off a pack.
Introduction: why choosing skincare is no longer guesswork

The classic formula-development workflow looks like this: a formulator reaches for actives that "work for everyone" — Niacinamide for post-acne marks, Retinol for wrinkles, Hyaluronic Acid for hydration — and assembles a cream from a template proven on previous clients. Then comes the trial-and-error stage: testing on a panel of 15-20 volunteers, two or three formula iterations, an 8-12 week wait for results. If the skin reacts unexpectedly — irritation, comedogenicity, no visible effect — the cycle starts over. A single finished product takes anywhere from 6 months to a year and a half, and the testing budget for a large line easily exceeds the cost of developing the formula itself.
Those workflow figures are an editorial description of common industry practice, not a measured statistic — we flag that here because the rest of this article is strict about the difference. Throughout, every number is labelled by the kind of document it comes from: a peer-reviewed study, a regulation, a supplier technical page (authoritative for that supplier's raw material and nothing else), a company claim (marketing until someone publishes a method), or an editorial observation from teaching practice.
The problem isn't the formulators' expertise — it's the logic of the approach itself: a one-size-fits-all lineup of actives averages out a skin response that is, by its very physiology, extremely heterogeneous. Barrier function, microbiome, sebum secretion rate, desquamation speed and sensitivity to irritants vary so much between people with the "same skin type" on paper that the identical 2% Salicylic Acid visibly reduces inflammatory lesions in two weeks for one client and triggers dryness and flaking within three days for another.
Where the traditional matching model breaks down
The trial-and-error approach rests on three assumptions, each statistically weak:
- Skin type as a category — dividing skin into "dry/oily/combination" ignores dozens of biomarkers (transepidermal water loss, surface pH, pilosebaceous follicle density) that actually determine the response to an active. This is not a rhetorical point: when a multi-centre survey built 2,737 species-level metagenome-assembled genomes from facial skin and let the data decide the groupings, it recovered two reproducible "cutotypes" driven chiefly by age, alongside measurable variation in moisture, sebum, gloss, pH, elasticity and sensitivity (peer-reviewed study, Microbiome, 2024). Neither cutotype corresponds to the retail categories;
- Transferable efficacy between people — an active's average efficacy in a clinical study doesn't guarantee the same result on a specific person's skin. Concretely: for facial hyperpigmentation, split-face trials collected in a 2021 mechanistic review found that a 5 % nicotinamide moisturiser outperformed vehicle at 4 and 8 weeks while a 2 % moisturiser showed no statistically significant effect against vehicle at all (peer-reviewed review, Antioxidants, 2021) — and the same review records the spot-reduction advantage fading back to vehicle level by week 42 after treatment stopped;
- A static formula — the recipe is fixed once, while skin changes with the season, hormonal cycle and external factors — and the formula doesn't.
As a result, the cosmetics industry has spent years compensating for this uncertainty with SKU count: a line of 8-10 creams "for different needs" instead of a single formula genuinely tailored to the specific case.
What neural networks change
The thesis of this article is simple: machine learning turns active selection from categorical guesswork into a data-driven prediction task. Instead of asking "which active usually helps with acne," the model answers "which combination of actives will give the maximum effect for a specific skin profile — with a known sebum level, barrier function and response to previous products." Three classes of models are used for this:
- predictive models based on clinical and consumer data (regression, gradient boosting) that estimate the probability of irritation or efficacy of an active for a given profile;
- generative models that propose combinations of active concentrations within permitted technological ranges;
- computer vision for analyzing skin photos and extracting features (pore size, redness, pigmentation spots) that become input data for the first two model types.
This doesn't replace formula chemistry or override biochemical mechanisms — Retinol still needs time for the stratum corneum to adapt, and Ascorbic Acid is still unstable and poorly absorbed above pH 3.5, a threshold established by percutaneous absorption work (Pinnell et al., Dermatologic Surgery, 2001, PMID 11207686). What changes isn't the nature of the actives themselves, but the accuracy of hitting the right concentration and combination on the first try.
In the following sections we'll break down what data a model needs for accurate active selection, how algorithms match an active to a formula's task (acne, pigmentation, barrier, aging), and where the technology is already being applied in real product development — with examples accessible without a costly corporate R&D platform.
How a neural network "sees" skin: from photo to digital profile

Before an algorithm can propose a single formula, it has to turn a facial photo into a set of numbers. This process is called feature extraction, and its quality determines whether a client receives a recommendation based on the real state of her skin — rather than an averaged "one for everyone" template.
Computer vision processes a skin image through several sequential stages, each solving its own narrow task.
Stage 1: facial-zone segmentation
The first thing the system does is map the face into anatomical zones: T-zone, cheeks, periorbital area, nasolabial folds, chin. This matters because skin behaves differently across zones even on the same person: the T-zone may show oiliness with enlarged pores while the cheeks show dehydration and flaking. Without segmentation, the model would average this data and produce a contradictory recommendation.
Segmentation is performed using convolutional neural networks (CNNs) trained on annotated dermatological image databases. The model finds facial key points (landmarks) — typically 68-468 points, depending on the architecture — and builds a zone mask from them. Landmark counts are a documented property of the published architectures themselves. Segmentation accuracy is a different matter, and this is our first correction to the usual telling of this story: figures in the 95–98 % range circulate widely in trade write-ups without a traceable primary study behind them, so this article does not assert one. What can be stated from published work is the direction of the problem, not a headline percentage — and the next section on limitations shows how large the direction can be.
Stage 2: detection of specific features

Within each zone, the algorithm looks for specific visual patterns that are converted into measurable metrics:
- Pores — detected as local roundish dark spots of a specific diameter (typically 50-200 microns after pixel conversion); density per cm² and average size are calculated
- Redness (erythema) — analysis of the RGB red channel and its deviation from the baseline skin tone; an erythema-intensity map is built on a scale
- Sebum and shine — determined via analysis of specular reflections (glare highlights) in the image; the higher the density and brightness of the highlights, the higher the estimated sebum secretion level
- Wrinkles and creases — detection of linear structures using Gabor- or Frangi-type filters that pick out elongated dark lines against a lighter texture background
- Pigmentation and post-acne marks — clustering of irregularly shaped dark spots, distinguished from pores by size and contour
Each of these features is not just flagged as "present/absent" — it receives a numeric value: a density index, a severity degree, the affected area as a percentage of the total segment area.
That these extracted features carry real signal is not merely assumed. In the largest published deployment of this kind — an AI analysis of smartphone selfies from 1,156,703 adult Chinese women, combined with a lifestyle questionnaire — the machine-scored features of blackheads, pore severity, dark circles and skin roughness were each positively associated with acne severity, and the study recovered a clear age curve: severity declining from age 25, a trough at 40–44, then a gradual rise (peer-reviewed study, Skin Research and Technology, 2024). Note the shape of that evidence: it establishes that the features are informative across a population. It does not establish that any one selfie yields a reliable individual diagnosis.
Stage 3: normalization and conversion into a feature vector
Raw data from a photo is heavily dependent on shooting conditions: lighting, camera angle, screen color temperature. So before the data is fed into the active-selection model, it goes through normalization — bringing it to a single scale invariant to shooting conditions. This is often done via calibration against a reference color card (color checker) or against reference facial zones with known characteristics (for example, the eye sclera as a white reference).
After normalization, all visual features are assembled into a single feature vector — an ordered set of numbers where each position corresponds to a specific parameter: pore density in the T-zone, erythema index on the cheeks, nasolabial wrinkle depth, and so on. Such a vector typically contains 30 to 150 parameters, depending on the model's level of detail.
| Visual feature | Detection method | Resulting profile parameter |
|---|---|---|
| Pores | Blob detection, local contrast | Density per cm², average diameter |
| Redness | RGB/LAB analysis, erythema map | Inflammation index 0-100 |
| Sebum | Specular-highlight analysis | Oiliness index by zone |
| Wrinkles | Gabor/Frangi filters | Line density and depth |
| Pigmentation | Dark-spot clustering | Spot area and contrast |
It is this numeric vector — not the photo itself — that is passed on to the active-ingredient selection model. The model no longer "looks" at the face; it works with structured data, comparing the client's profile to a training database where similar feature combinations were linked to specific formulas and their efficacy. That matching process is covered in the next section.
Compatibility chemistry: why not all actives can be mixed
Even an active perfectly matched to a skin type may fail to work — or sometimes cause harm — if the formula pairs it with "neighbors" it conflicts with biochemically. Before AI proposes an ingredient combination, it has to model dozens of hidden interactions: competition for the same receptors, shifts in pH, oxidative instability, competitive adsorption on the lipid barrier. This is "compatibility chemistry" — the invisible layer of constraints that determines what formulation space is even available.
Retinol and AHA: a conflict of renewal mechanisms
Retinol works by binding to nuclear RAR/RXR receptors, triggering expression of genes responsible for keratinocyte proliferation. Alpha Hydroxy Acids (glycolic, lactic acid) work differently — they break ionic bonds between corneocytes in the stratum corneum, accelerating mechanical exfoliation. The problem isn't that these mechanisms "contradict" each other — they add up, and that's exactly what's dangerous: two overlapping cycles of accelerated renewal leave the barrier less time to recover, with rising transepidermal water loss (TEWL) and erythema as the expected consequence.
Here we owe readers a correction rather than a reassurance. Earlier versions of this piece attached a specific threshold — retinol at 0.3 % together with AHA above 5 % — to a named citation we have been unable to locate in the literature. The mechanism above is well established; that exact numerical pairing, as a published finding, is not, and we have removed the attribution rather than soften it with "typically". What a formulator can act on instead is measurable: TEWL is the standard readout for this kind of barrier stress, and the method itself is standardised in the revised EEMCO guidance for in vivo measurement of water in the skin (Skin Research and Technology, 2018, PMID 29923639). If you combine renewal actives, the defensible move is to measure TEWL on a panel in your own system rather than to inherit a number from an article.
An additional layer of complexity is pH. Retinol is stable in the pH 5.5-6.5 range, while AHA efficacy drops sharply above pH 4. A formula trying to "straddle" two environments at once usually loses stability for both actives: the retinol oxidizes faster, and the acid doesn't deprotonate to the degree needed to penetrate the stratum corneum.
Vitamin C and niacinamide: a myth that has outlived its shelf life
A classic worry: L-Ascorbic Acid and Niacinamide supposedly form niacin and cause flushing. This is true only when the mixture is heated above 60°C in an aqueous medium during prolonged incubation — a condition irrelevant to a finished cosmetic emulsion at room temperature. The mechanistic review by Wohlrab and Kreft (Skin Pharmacology and Physiology, 2014, PMID 24993939) sets out the pharmacology of topical nicotinamide and gives no basis for expecting clinically significant transamidation-driven irritation under ordinary conditions of use. Still, an AI model needs to account not for the reaction itself, but for the conditions that accelerate it: the manufacturing process, storage temperature, the final pH of the finished formula. If a cream is formulated with L-Ascorbic Acid at pH 3.5 (the stability and absorption optimum for this form of vitamin C, per Pinnell et al., 2001), adding niacinamide at a high concentration creates an uneven buffer system where part of the molecules migrate into a suboptimal pH zone.
| Active | Optimal pH | Conflicting partner | Conflict mechanism |
|---|---|---|---|
| Retinol | 5.5-6.5 | AHA/BHA >5% | Cumulative increase in desquamation, rising TEWL |
| L-Ascorbic Acid | 3.0-3.5 | Niacinamide (high conc.) | Buffer shift, potential transamidation under heat |
| Sodium Ascorbyl Phosphate | 6.0-7.0 | Acid exfoliants | Hydrolysis of the phosphate group, loss of activity |
| Benzoyl Peroxide | 4.5-5.5 | Retinol, vitamin C | Oxidative degradation of both partners |
How AI models these constraints
A generative model doesn't "invent" a formula freely — it operates within a constrained optimization space, where each active ingredient is described not just by its usage rate but also by a vector of physicochemical parameters: stable pH range, decomposition temperature, redox potential, known incompatibility pairs from the literature and patent databases. In practice this is implemented as a penalty-function system: if the algorithm tries to combine retinol and AHA above a threshold concentration, the overall "formula score" drops, and the model searches for an alternative combination — for example, encapsulated retinol or PHA instead of glycolic acid, which is milder in its irritation mechanism.
A working example of this pattern from outside skincare recommendation shows what a defensible version looks like. Two ligand-based machine learning models — one classifying activity, one estimating pIC50 — screened natural-product and drug libraries for tyrosinase inhibitors; molecular docking re-ranked the survivors; and the three top candidates were then tested in vitro, all inhibiting mushroom tyrosinase more strongly than arbutin, with better transdermal permeation than commercial arbutin formulations (peer-reviewed study, ACS Omega, 2025). The constraint layer there is the wet lab. That is the standard against which a cosmetic recommender's internal "compatibility score" should be judged — and almost none of them publish one.
That's exactly why the solution space AI actually sees is much narrower than it first appears: out of thousands of theoretical active combinations, only a minority remain chemically compatible once pH, temperature and oxidative stability are taken into account. The "15-20 %" figure once quoted here was an editorial estimate, not a measurement, and is presented as such: the surviving share depends entirely on how many actives are in the library and how strict the penalty thresholds are. More on this in the article on cosmetic formula stability.
Recommendation-engine architecture: from user data to a finished formula
Under the hood of any "AI picks your cream for you" service is not one model but a pipeline of several data-processing stages. Each stage solves its own narrow task, and only their sequential chaining turns disparate inputs — a photo, a questionnaire, complaints — into a specific list of actives with usage rates. Let's break down this pipeline link by link.
The input layer: what the model receives
The recommendation system works with three types of data, which are merged into a single feature vector before being fed into the model:
| Data source | What's extracted | Format for the model |
|---|---|---|
| User questionnaire | self-reported skin type, complaints, climate, hormonal status, current routine | categorical and binary features |
| Facial photo | texture, pore visibility, redness, pigmentation, UV damage (with multispectral imaging) | embedding from a convolutional network (typically 128-512 features) |
| Target task | "calm inflammation," "even out tone," "anti-aging," "barrier repair" | one-hot goal vector |
An important nuance: questionnaire data and the photo embedding live on different numeric scales, so before merging them normalization is applied (usually z-score or min-max) — otherwise one data source starts dominating the other purely due to the range of its values, not its actual importance.
Model block one: profile classification
The first stage solves a multi-class or multi-label classification task: the system doesn't assign the user's skin to a single type (a simplification from 1990s skincare questionnaires), but to a combination of states — for example, "seborrheic T-zone + peripheral dehydration + post-acne pigmentation." Gradient boosting (XGBoost, CatBoost) on tabular features is often used for this, or a small fully connected network if the input is a merged photo-plus-questionnaire embedding.
How accurate is that classification against a dermatologist's assessment? This is the second place where we have removed a number rather than hedge it: the "78–89 %" range previously cited here traces to no locatable publication. The honest state of the evidence is that consumer skin-classification accuracy is largely reported by the companies selling the classifiers, and that where independent evaluation has been done on dermatological imaging, the results were sobering rather than reassuring — see the limitations section below. Any vendor quoting a single accuracy figure without naming the test set, the reference standard and the population it was measured on is making a company claim, not reporting a result.
Block two: the recommendation engine for active selection
After profile classification, a second model kicks in — the actual recommendation engine, which works on logic close to e-commerce product recommendations, but with strict chemical constraints. Two approaches are used here, often in a hybrid:
- Content-based filtering — matching the skin profile against a knowledge base of actives: each ingredient is pre-tagged with target conditions, mechanism of action, effective concentration range, and known incompatibilities.
- Collaborative filtering — training on historical data about how similar users responded to similar formulas (only works with a sufficiently large base of reviews and repeat purchases — tens of thousands of profiles).
The content-based component answers "what theoretically fits"; collaborative filtering answers "what actually worked for similar people." Without the first, the system recommends unsafe combinations; without the second, it loses sensitivity to real skin-response patterns that don't always match the theory.
There is a structural problem with the knowledge base that no amount of model architecture fixes, and it is legal rather than technical. If the "effective concentration range" tags are scraped from marketed products' INCI lists, they are fiction: under Article 19 of Regulation 1223/2009 as retained in UK law, the ingredient list runs in descending order of weight only above 1 %, everything below 1 % may be listed in any order, and no concentration is ever declared on pack. A recommender whose concentrations come from labels has inferred them; a recommender whose concentrations come from clinical papers and supplier technical documents has read them. Those are not the same system, and the difference is invisible to the user.
The constraint layer: why the model can't just maximize efficacy
The final list of actives isn't chosen purely on "what will give the maximum effect" — sitting downstream of the recommendation model is a rule-based constraint layer that filters combinations by chemical compatibility, the formula's total acidity, and the maximum allowed concentration of irritating actives for a given sensitivity level. This is exactly the layer that stops the system from simultaneously assigning 10% Niacinamide and a high dose of pure acids to skin showing signs of a compromised barrier.
The output layer: what the user sees
The end result of the pipeline isn't a text description — it's a structured list with usage rates and application order. Percentages below are active matter; a commercial raw material supplying that active at, say, 50 % activity would be dosed at twice these figures as supplied (w/w):
| Active | Concentration | Role in the formula |
|---|---|---|
| Niacinamide | 4% | barrier reinforcement, sebum regulation |
| Sodium Hyaluronate | 1% | hydration, surface smoothing |
| Bakuchiol | 0.5% | anti-aging effect without retinoid irritation |
Each of those three rows has a different evidential status, and a formulator should hold them differently. The 4 % niacinamide row is supported by split-face trial evidence for wrinkles, pores and unevenness at 8–12 weeks, collected in the 2021 Antioxidants review and originating with Bissett's work on topical niacinamide in ageing facial skin (International Journal of Cosmetic Science, 2004, PMID 18492135). The 0.5 % bakuchiol row rests on a single well-known head-to-head: a randomised, double-blind twelve-week comparison of 0.5 % bakuchiol against 0.5 % retinol found comparable improvement in wrinkle area and pigmentation with less scaling and stinging in the bakuchiol arm (Dhaliwal et al., British Journal of Dermatology, 2019, PMID 29947134) — one trial, not a body of evidence. The sodium hyaluronate row is conventional practice rather than a cited threshold. An algorithm that outputs all three with the same confidence formatting is hiding that difference from you.
This level of detail is a direct consequence of the architecture: the system doesn't "guess a product" — it assembles a formula from atomic decisions, each of which has passed through classification, recommendation, and a compatibility filter. More on how these percentages map to real lab usage rates is covered in the piece on active-ingredient dosage rates.
Skin types through the algorithm's eyes: oily, dry, combination — and what the model sees
When a user checks "I have oily skin" on a questionnaire, she's describing a subjective feeling — shine by midday, clogged pores, foundation "sliding off" by evening. The algorithm works with a different layer of reality: it operates with measurable physiological parameters that don't depend on mood, the weather outside, or how much coffee was drunk that morning. The gap between these two pictures of reality is exactly why neural-network skin-type classification can be more informative than self-reporting — provided the instrument behind it has actually been validated.
Three markers the model looks at
Most algorithmic skin classifiers are built on three physiological variables that dermatologists have measured instrumentally since the 1980s, and which neural networks have learned to estimate from photos and indirect signs:
- Sebum rate — the rate and volume of sebum production by sebocytes. Measured clinically with a sebometer (in μg/cm²); from a photo the algorithm estimates it indirectly — via specular-glare patterns in the T-zone, the size and density of visible pores, and local pixel texture around the nose and forehead.
- Transepidermal water loss, TEWL — the rate of water evaporation through the stratum corneum, a key indicator of lipid barrier integrity. High TEWL signals dryness and heightened skin reactivity even when the skin looks normal visually. From a photo this is read via microtexture: flaking, uneven relief, areas with a matte, "tight" surface. The instrumental method and its confounders — ambient humidity, air movement, skin temperature, acclimatisation time — are standardised in the revised EEMCO guidance (2018, PMID 29923639), which is worth reading before quoting anyone's TEWL number, including your own.
- Elasticity and firmness — the skin's ability to return to its original position after mechanical deformation, linked to the state of the dermis's collagen and elastin matrix. Measured instrumentally with a cutometer; the algorithm estimates it from wrinkles, microrelief and contour sagging over time (if a series of photos or video is available).
The combination of these three parameters produces not a single label of "oily/dry/combination" but a vector of values — for example, sebum rate 68% (high), TEWL 12 g/m²/h (typically reported in the single digits for undamaged facial skin under controlled conditions), elasticity 0.75 (reduced). Classification is built on this vector, not on a visual "looks like oily skin." Absolute TEWL cut-offs vary between devices and laboratories, which is precisely why the EEMCO guidance insists that measurements be compared within a study rather than across them.
Why self-assessment systematically gets it wrong
The gap between what people think about their own skin and what instruments show is real, but here too we have pulled a number rather than dress it up: the frequently repeated "fewer than 60 % of self-diagnoses match sebometer readings" has no source we could verify, and it is not stated as fact in this revision. What is supportable is that the underlying properties genuinely vary in ways self-report cannot track — the multi-centre facial survey cited earlier measured moisture, sebum, gloss, pH, elasticity and sensitivity as continuous variables and found age to be their dominant driver, with no clean boundary reproducing the retail categories (Microbiome, 2024). The mechanisms that make self-assessment unreliable are well understood:
- Sebum output fluctuates across the day and with ambient temperature, so an "eyeballed" assessment in the morning and evening gives different answers.
- Dryness is often confused with dehydration (a lack of moisture in the stratum corneum rather than lipids) — these are two different mechanisms requiring different actives.
- Combination skin is, in our teaching experience, the most frequently misdiagnosed category: people either overestimate T-zone oiliness or don't notice cheek dryness because their attention is drawn to visible problems (pores, shine). This is an editorial observation from student intake, not a published statistic.
- Seasonality and "masking" makeup: foundation or a mattifying primer visually normalize the skin in the moment, distorting perception of its baseline state.
Combination skin: why it isn't "two types in one"
The algorithm treats combination skin not as a mix of oily and dry, but as zonal heterogeneity within the same set of parameters. The model divides the face into segments (T-zone, cheeks, periorbital area) and calculates its own sebum/TEWL/elasticity vector for each. The resulting profile isn't an averaged value but a map of differences: the T-zone may show a high sebum rate at low TEWL, while the cheeks show low sebum at markedly higher TEWL. This is critical for formulation: a serum with a single Niacinamide concentration for the whole face performs worse than a zone-specific active distribution, a point covered in detail in the piece on ingredient compatibility (/blog/himiya-sovmestimosti-aktivov).
Practical takeaway for formulation: the accuracy of skin-type classification directly determines the permissible active concentration range. An error in TEWL estimation could lead to a formula with 2% Salicylic Acid being recommended for skin with an already compromised barrier — and instead of clearer pores, the user ends up with irritation and heavier flaking.
The formula's task: how AI tells "anti-acne" apart from "for radiance"
Skin type sets the formula's baseline constraints, but it's the target task that determines which actives make it into the composition first. For the algorithm this isn't a text label like "acne" or "radiance" — it's a vector of clinical endpoints, a set of measurable parameters against which an ingredient has proven its efficacy in studies. The model matches the user's request against this vector, not against a problem name.
From a query to a clinical vector
When a user writes "I want to get rid of breakouts," the system's NLP module translates the request into a combination of features: follicular hyperkeratinization, Cutibacterium acnes colonization, local inflammation, excess sebum production. A "for radiance" request decomposes differently: uneven stratum corneum, slowed cell turnover, micro-dullness from collagen glycation. This difference in request decomposition is the first level at which tasks are distinguished.
From there, each feature searches for a match in the actives database, where every ingredient is tagged not by marketing claim but by mechanism of action confirmed in studies. Concentrations in the table below are active matter:
| Task | Key mechanism | Typical actives | Concentration range |
|---|---|---|---|
| Anti-acne | Sebum regulation + antimicrobial action | Salicylic Acid, Azelaic Acid, Niacinamide | 0.5-2% / 10-20% / 4-5% |
| For radiance | Accelerated desquamation + antioxidant protection | Glycolic Acid, Sodium Ascorbyl Phosphate, Tranexamic Acid | 5-10% / 3-5% / 2-3% |
| Anti-aging | Stimulation of collagen synthesis | Retinol, peptide complexes, Bakuchiol | 0.3-1% / 2-5% / 0.5-2% |
| Barrier repair | Lipid replenishment | Ceramide NP, Cholesterol, Squalane | 1-3% / 0.5-1% / up to 10% |
Importantly, the same ingredient can appear across several tasks with a different weighting function. Niacinamide works on acne (sebum regulation), pigmentation (inhibiting melanosome transfer) and barrier repair (stimulating ceramide synthesis) alike — the mechanisms are set out in Bissett et al. (2004) and in the later mechanistic review already cited. The algorithm doesn't pick an ingredient "for a task" — it sums its mechanisms' contribution across all of a user's active target vectors at once.
How clinical data calibrates the weights
Every "active → effect" link in the database carries not a binary value (works/doesn't work) but a numeric efficacy coefficient derived from trials and reviews. For Azelaic Acid, the authoritative summary of its multi-target mechanism — anti-inflammatory, antioxidative, bactericidal including against antibiotic-resistant strains, and anti-comedogenic — is the review by Sieber and Hegel (Skin Pharmacology and Physiology, 2014, PMID 24280644). Two caveats belong with that citation and are the kind an algorithm never surfaces: it is a narrative review rather than a meta-analysis, and its authors were employed by the manufacturer of an azelaic acid product at the time. Head-to-head efficacy figures for 15 % azelaic acid gel come from separate randomised trials (for example the investigator-blind parallel-group study in the Journal of the European Academy of Dermatology and Venereology, 2015, PMID 25399481), not from that review — a distinction that matters, because a review and three retellings of the same trial are one piece of evidence, not four.
The niacinamide case makes the same point more sharply, and it is where the earlier version of this article was simply wrong. It stated that a clinically noticeable effect on sebum begins at 4–5 %. The study it cited says the opposite: Draelos et al. tested 2 % niacinamide and found significantly lowered sebum excretion rate after 2 and 4 weeks in a placebo-controlled trial of 100 Japanese subjects, plus significantly reduced casual sebum levels after 6 weeks in a 30-subject split-face study in Caucasian subjects (Journal of Cosmetic and Laser Therapy, 2006, PMID 16766489). The 4–5 % threshold belongs to a different endpoint entirely — facial hyperpigmentation, where 5 % beat vehicle and 2 % did not (Antioxidants, 2021). The correct generalisation is therefore not a number but a rule: an active's effective concentration is defined jointly with its endpoint, and a single "recommended %" for an ingredient is a category error. This is exactly the kind of collapse a recommendation engine performs when it stores one usage-rate field per ingredient.
A second calibration source is aggregated feedback. When thousands of users with similar digital skin profiles report changes in specific parameters (oily shine, papule count, skin tone) after 4, 8 and 12 weeks of use, the system recalculates recommendation weights. This works like a Bayesian update: clinical data sets the prior efficacy distribution, and behavioral feedback provides a posterior correction for the real user population — not just for laboratory-study conditions.
Conflicting tasks within a single formula
Complexity arises when a user states two goals at once — "clear acne and add radiance." The vectors partially overlap (both goals need desquamation) but partially conflict: aggressive sebum regulation can worsen the dryness and dullness the user is trying to fix. The algorithm resolves this by prioritizing based on condition severity — active inflammation gets a higher weight than an aesthetic goal, while supporting actives (antioxidants, for instance) are introduced at a supportive rather than a high concentration. More on how the system resolves such overlaps at the ingredient-compatibility level is covered in the piece on active-compatibility chemistry.
Limitations and risks: when the algorithm gets it wrong
A recommendation system is only as good as the data it was trained on. This rule sounds banal, but it explains most of the systematic errors in AI active selection. The model doesn't "understand" skin — it finds statistical patterns in a sample, and if the sample is skewed, so is the result.
Bias by phototype and ethnicity
Most open dermatological datasets used to train computer-vision models for skin analysis have historically consisted predominantly of images of lighter phototypes (I-III on the Fitzpatrick scale). Post-inflammatory hyperpigmentation, common in phototypes IV-VI, is detected worse by such models: the algorithm confuses it with an inflammatory lesion or fails to flag it as a distinct issue altogether. Adamson and Smith set out the mechanism in JAMA Dermatology (2018, PMID 30073260): a lack of diversity in dermatological image datasets systematically lowers accuracy specifically on medium and dark skin tones.
That argument was subsequently measured rather than asserted. Researchers built the Diverse Dermatology Images (DDI) dataset — the first publicly available, expert-curated and pathologically confirmed image set with diverse skin tones — and showed that state-of-the-art dermatology AI models performed substantially worse on it, particularly on dark skin tones and uncommon diseases; notably, the dermatologists who label such datasets also performed worse on those same images, which is how the bias gets baked in. Fine-tuning the models on DDI images closed the light/dark performance gap (peer-reviewed study, Daneshjou et al., Science Advances, 2022). For active selection this creates a concrete risk: the system may fail to recommend brightening ingredients like Tranexamic Acid or Alpha Arbutin where they're objectively needed, simply because it didn't recognize the pigmentation pattern.
Overfitting to marketing data
Some recommendation services are trained not on clinical data but on product-listing copy, reviews, and brand marketing descriptions. In such a dataset, words like "hydrates," "minimizes pores," "evens out tone" appear far more often than actual TEWL (transepidermal water loss) or sebumetry measurements. The model starts overfitting to the seller's rhetoric rather than to the ingredient's biochemistry. And it has no corrective available on the label itself: as noted above, Article 19 of the retained Regulation 1223/2009 means the pack carries an ordered list above 1 %, an unordered list below it, and no numbers at all. A system that "learns concentrations from products" is learning from text that does not contain them.
Lack of medical history
None of the mass-market consumer skin-photo-analysis systems has access to a history of allergic reactions, systemic medications (isotretinoin, anticoagulants, hormone therapy), or coexisting dermatoses — rosacea, perioral dermatitis, atopy. The algorithm sees texture and color, but doesn't "know" that a user is taking warfarin, or that a high concentration of Retinol combined with acid exfoliants will raise irritation risk on top of a fine vascular network. No visual analysis replaces a patch test and history-taking.
Where AI does have a documented role in cosmetic safety, it looks quite different from photo triage: an explainable deep-learning framework combining temporal convolutional and LSTM layers with protein language-model embeddings was built to flag allergenic motifs in peptide sequences before those peptides enter a formulation, with Anchor, LIME and SHAP exposing which motifs drove each call so a human can audit the reasoning (peer-reviewed study, Journal of Proteome Research, 2026). That is AI applied where the input is a known sequence and the output is auditable — the opposite situation from inferring a person's medical history from a selfie.
| Limitation type | Error mechanism | Consequence for the formula |
|---|---|---|
| Phototype data bias | Lack of dark-skin images in the training sample | Missed PIH, incorrect brightening-active selection |
| Overfitting to marketing | Training on ad copy instead of clinical data | Under- or overestimating the effective concentration |
| No medical history | No access to allergies and medications | Risk of interaction with systemic drugs |
| No dermatological validation | Model not tested on real patients with pathology | False confidence in the formula's safety |
Why dermatological validation is needed
A model that shows high accuracy on an internal test set isn't automatically clinically valid. The gap between accuracy on a dataset and real diagnostic value is a standard ML-in-medicine problem, and the DDI work above is the cleanest cosmetic-adjacent demonstration of it: models that performed well on their original benchmarks degraded on an independent, pathology-confirmed set assembled specifically to include the populations the benchmarks under-represented. For cosmetic recommendation systems, this means a claimed "90% accuracy" almost always refers to a narrow sample and doesn't guarantee the same result on a specific user's face. Ask three questions of any such figure: which dataset, which reference standard, and which population.
Practical takeaway for anyone designing or using such systems: an AI recommendation is a formula draft that requires verification on real skin, a patch test, and — in borderline cases — a dermatologist consultation. Automation speeds up active selection, but it doesn't remove responsibility for the final decision from the person composing the formula.
The future of personalization: from recommendation to made-to-order production
Everything described above — skin recognition, active-compatibility checks, formula-task matching — operates within the existing assortment: the algorithm picks a finished product or, at most, suggests "mixing serum A with cream B." The industry's next step is to remove the intermediary of a ready-made product line altogether. Instead of a recommendation, the system is handed the authority to generate a composition from scratch for a specific digital skin profile — and then manufacture exactly that formula as a single unit.
This is a fundamentally different business architecture: not "10,000 identical jars" but a micro-batch for one person — a batch of 1 to 50 ml, assembled to parameters that no other client shares. Technically, this became possible thanks to three converging directions: generative models capable of proposing not a finished product but a vector of concentrations; modular production lines with dosing units that blend actives in real time; and 3D printing of cosmetic forms, where the printer serves as the final assembly step.
Generative selection instead of classification
Current recommendation systems solve a classification task: "this skin profile fits product X from the catalog." A generative model solves an optimization task: "which combination of 15-20 permitted actives, in what proportions, gives the maximum expected effect at zero incompatibility risk." The input is the same digital skin profile (texture, sebumetry, hydration, response history), but the output isn't a product name — it's a numeric vector: for example, Niacinamide 4%, Sodium Hyaluronate 1%, Panthenol 3%, base to 100%.
The claim that this has been validated against human formulators is one we can no longer support in the form it previously took here: the citation attached to it is not locatable, and we are not going to launder an untraceable result by calling it "reportedly". What is publicly demonstrable today is the narrower pattern described earlier — machine learning proposing ranked candidates that a laboratory then confirms or rejects, as in the tyrosinase-inhibitor screen. Everything beyond that, in cosmetic formulation specifically, currently sits in corporate pipelines and press releases rather than in the literature. Treating it as established is how a technology roadmap turns into marketing.
3D printing: from concept to shelf
3D printing of cosmetics is no longer a futuristic metaphor — extrusion printers for semi-solid textures (balms, masks, sticks) handle a viscosity of 5,000-50,000 mPa·s and allow different actives to be layered within a single item: for example, a zone with an elevated Salicylic Acid concentration across the T-zone of a mask, and a zone with Ceramide NP around the face's periphery. Printing temperature ranges are limited by the thermal sensitivity of the actives: peptides and vitamin C in the form of Ascorbic Acid tolerate extrusion heat poorly, which pushes manufacturers toward stable derivatives — Sodium Ascorbyl Phosphate or Ascorbyl Glucoside. The specific ceiling temperature depends on the derivative, the pH and the residence time at temperature, and belongs in the supplier's technical documentation for the grade you are actually buying, not in a general article.
| Parameter | Mass production | Personalized micro-batch |
|---|---|---|
| Batch volume | 1,000-50,000 L | 1-50 ml |
| Order-to-ready time | weeks (logistics + warehouse sale) | minutes-hours (local assembly) |
| Formula source | fixed R&D-department formula | generative model + user profile |
| Stability control | centralized, at the manufacturing stage | distributed, requires local AI monitoring |
What still has to mature in biotech
Generative personalization is held back less by model power and more by the constraints of chemistry and biotech. Three bottlenecks:
- Stability at small volumes. A micro-batch without a preservative system designed for industrial scale spoils faster — new preservative-selection protocols are needed for a 10-50 ml volume, given a 2-4 week usage window. What that looks like concretely, at the level of a real document: INOLEX specifies its SpectraStat GHL Natural blend (Propanediol, Caprylhydroxamic Acid, Glyceryl Heptanoate) at 1.0–2.0 % as supplied (w/w), narrowing to 1.0–1.5 % in an O/W emulsion, added to the water phase below 80 °C, with pass results quoted against EP-A, EP-B, USP 51 and PCPC challenge tests (supplier technical page). Note what a personalised micro-batch destroys: those pass results were obtained in a specific base. A formula recomposed per customer has never been challenge-tested in the composition it actually ships in.
- Biosensors for input data. A generative formula's accuracy depends directly on the accuracy of the skin profile: wearable biosensor development (real-time TEWL, pH, microbiome measurement) needs to catch up with the algorithms' ambitions.
- Regulatory framework. Cosmetics legislation — Regulation 1223/2009 as retained in UK law for the GB market, and the EU version with its CosIng annexes for products sold into the EU — was written for serial production with a fixed composition, with a responsible person, a product information file and a safety assessment attached to that composition. There is still no settled route by which a formula that changes with every customer's unique batch satisfies that structure.
A realistic horizon for the mass shift from recommendation to made-to-order production is, in our editorial judgement, several years rather than several quarters, and the first to adopt it won't be mass-market brands but niche dermatology clinics and premium labs, where the cost of a single micro-batch is justified by medical precision. That is a forecast, clearly labelled as one. For specialists learning to formulate by hand today, this trend isn't a threat — it's a shift in role: from creating one formula for thousands of people to setting the rules by which an algorithm creates thousands of formulas for one person.
Conclusion: trust the algorithm or trust your skin
Across nine sections we've traced the path from a skin photo to a production line capable of assembling a serum around a specific data set. But the ultimate question is always the same: who decides what ends up on a face — the model or the person. The answer isn't binary. The algorithm is precise about what's measurable: Niacinamide concentration, Retinol-and-AHA compatibility, the pH window for Ascorbic Acid stability. It's blind to what requires context — hormonal background, pregnancy, isotretinoin use, a person's psychological relationship to a product's texture and scent.
It is worth stating plainly what this revision found when it went back to the sources. Of the numeric claims originally carried in this article, several traced to citations that could not be located at all, and one — the assertion that niacinamide's effect on sebum begins at 4–5 % — was contradicted by the very study it cited, which used 2 % and worked. That is not a criticism of AI; it is a demonstration of the failure mode AI inherits, since a model trained on articles like the earlier version of this one would have learned the wrong number with perfect fluency. The editorial conclusion is that a single universal usage rate per active cannot be established, and should not be stated — concentration is meaningful only jointly with an endpoint, a system and a test method.
What the algorithm does better than a person
A neural network never tires of cross-checking dozens of parameters at once and isn't subject to cognitive biases like "this ingredient is trendy, so it must be good." It holds the entire active-compatibility table in memory and recalculates the formula in seconds if even one input parameter changes — regional humidity, season, a new sebumetry reading. It's physically hard for a person to hold 40+ variables of INCI compatibility in mind during every single consultation.
| Task | Better handled by | Why |
|---|---|---|
| Calculating active % and pH compatibility | Algorithm | Strict math, no attention fatigue |
| Assessing visual skin markers from a photo | Algorithm (with caveats) | Standardization, no lighting subjectivity |
| Diagnosing a dermatosis requiring biopsy/history | Dermatologist | Requires data unavailable to the model: medical history, palpation, progression over time |
| Assessing sensory feel and subjective comfort of a formula | Cosmetologist + user | Aesthetics and tactile feel don't formalize into a dataset |
| Final safety decision during pregnancy, medication use | Physician | Legal and medical responsibility |
A practical model: algorithm as draft, human as editor
The working scheme that emerges from everything covered above looks like this: AI generates a composition hypothesis based on the digital skin profile and the stated task, and a specialist — a cosmetologist, dermatologist, or a sufficiently informed user — makes corrections the model couldn't account for. This isn't unlike how a GPS navigator works: it proposes a route based on map and traffic data, but the driver sees a closed road that isn't in the database and adjusts the route manually.
- Step 1. The algorithm analyzes the photo and questionnaire, outputs 2-3 formula options with key actives and their usage rates.
- Step 2. A cosmetologist or the user checks the option for obvious incompatibilities with the current routine — for example, prescription Tretinoin already in use, which the algorithm may not know about.
- Step 3. A patch test is run on the inner forearm for 48 hours, regardless of how "safe" the algorithm labeled the composition.
- Step 4. The skin's response is recorded and fed back into the system as a new data point for the next recommendation iteration.
A fifth step belongs in that list for anyone who formulates rather than merely consumes: check the provenance of every number the model hands you. Ask which document it came from — a regulation, a peer-reviewed study, the supplier's technical page for the exact grade you bought, or nobody at all. Where the model cannot say, the number does not go into the formula. Applied to this very article, that test removed four citations and reversed one conclusion.
What this means for those who formulate and sell cosmetics
For specialists developing formulas, AI-assisted selection isn't a threat to the profession — it's an expansion of the toolkit, comparable to the shift from manual emulsion calculations to software-based stability modeling. Understanding how a model interprets skin data and why it sometimes gets things wrong (see the section on limitations) is becoming part of a formulator's basic literacy — just like knowing emulsifier chemistry or the temperature regimes for Phase A and B.
The bottom-line position that holds up against both data and practice: the algorithm can be trusted with compatibility calculations and a first-pass concentration selection, but the final arbiter of safety and comfort remains a person — either a specialist with a medical background or the user herself, armed with an understanding of her own skin and the basic principles of active-ingredient chemistry. The technology shortens the guesswork, but it doesn't remove the need to listen to the skin's response here and now — it updates faster than any model can retrain.
Selecting actives is only the beginning — verifying compatibility and working concentrations comes next. The AI Chemist in the Walker Formulation Academy Club does this based on a verified INCI/CosIng database, not guesswork, and flags where actives conflict. If you want to master active selection systematically, the club offers a program with a live instructor.
Can an AI work out the percentages in a competitor's product from its INCI list?
No. Article 19 of Regulation 1223/2009 as retained in UK law requires descending order of weight only above 1 %; below that, ingredients may appear in any order, and no concentration is printed anywhere on the pack. Any percentage a model gives you for a marketed product is generated, not read.
Why does one source say 2 % niacinamide works and another says 5 %?
Because they are measuring different things. Draelos et al. (2006) found 2 % lowered sebum excretion; split-face trials for facial hyperpigmentation found 5 % beat vehicle while 2 % did not. An active's effective concentration is defined jointly with its endpoint, so a single "recommended %" per ingredient is meaningless without saying what effect you are buying.
Does a supplier's technical data sheet settle a dosage question?
It settles it for that supplier's raw material in that supplier's test system, and nothing further. A use level such as INOLEX's 1.0–2.0 % as supplied for SpectraStat GHL Natural comes with a stated phase, a stated process temperature and named challenge-test methods; transfer it to a different blend or a different base and the evidence does not travel with it.
Sources
- Facial adult female acne in China: an analysis based on artificial intelligence over one million. Skin Research and Technology, 2024 (PMID 38572573)
- Daneshjou R. et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Science Advances, 2022 (PMID 35960806)
- Mechanistic Basis and Clinical Evidence for the Applications of Nicotinamide (Niacinamide) to Control Skin Aging and Pigmentation. Antioxidants, 2021 (PMID 34439563)
- Integrated analysis of facial microbiome and skin physio-optical properties unveils cutotype-dependent aging. Microbiome, 2024 (PMID 39232827)
- Discovery of Potential Tyrosinase Inhibitors via Machine Learning and Molecular Docking with Experimental Validation. ACS Omega, 2025 (PMID 40918329)
- Deciphering Allergen Peptides for Dermatological and Cosmetic Applications with Explainable Artificial Intelligence. Journal of Proteome Research, 2026 (PMID 42345568)
- Regulation (EC) No 1223/2009 on cosmetic products, Article 19 (Labelling), as retained in UK law — legislation.gov.uk
- INOLEX, SpectraStat GHL Natural — product technical page (use levels as supplied, process conditions, challenge-test methods)
- Draelos Z.D., Matsubara A., Smiles K. The effect of 2% niacinamide on facial sebum production. Journal of Cosmetic and Laser Therapy, 2006 (PMID 16766489)
- Sieber M.A., Hegel J.K. Azelaic acid: properties and mode of action. Skin Pharmacology and Physiology, 2014 (PMID 24280644) — narrative review, manufacturer-affiliated authors
- Wohlrab J., Kreft D. Niacinamide — mechanisms of action and its topical use in dermatology. Skin Pharmacology and Physiology, 2014 (PMID 24993939)
- Bissett D.L. et al. Topical niacinamide reduces yellowing, wrinkling, red blotchiness and hyperpigmented spots in aging facial skin. International Journal of Cosmetic Science, 2004 (PMID 18492135)
- Dhaliwal S. et al. Prospective, randomized, double-blind assessment of topical bakuchiol and retinol for facial photoageing. British Journal of Dermatology, 2019 (PMID 29947134)
- Pinnell S.R. et al. Topical L-ascorbic acid: percutaneous absorption studies. Dermatologic Surgery, 2001 (PMID 11207686)
- Berardesca E. et al. The revised EEMCO guidance for the in vivo measurement of water in the skin. Skin Research and Technology, 2018 (PMID 29923639)
- Adamson A.S., Smith A. Machine learning and health care disparities in dermatology. JAMA Dermatology, 2018 (PMID 30073260)



