How to Break Down a Competitor's INCI List with AI: A Step-by-Step Guide

How to Break Down a Competitor's INCI List with AI: A Step-by-Step Guide

👩‍🔬 Walker Formulation Academy📅 19 August 2026⏱️ 24 min read

Why Analyze a Competitor's INCI List, and How AI Helps

Scientific illustration, flat design, of a cosmetic product label with an INCI ingredient list being scanned by a glowing AI neural network overlay, digital lines connecting text to a holographic structured formula diagram with layered circles representing water phase, oil phase, active phase, clean modern educational infographic style, cool blue and white color palette, no human faces
AI analyzes the INCI label and turns it into a structured formula

Every bottle on a mass-market shelf or in a niche boutique is more than marketing and packaging — it's an encoded formula, openly available to anyone in the form of an INCI list (International Nomenclature of Cosmetic Ingredients). That label listing is the only legal source from which a developer can extract the structure of someone else's formula without access to lab documentation. Analyzing a competitor's composition stopped being a niche pursuit of patent analysts long ago — it's now a routine task for brand managers, R&D specialists, and cosmetic chemists who want to understand the market faster than weeks of manual work would allow.

The INCI list follows a strict rule: ingredients are listed in descending order of concentration, while components present at less than 1% can appear in any order. That constraint is both a hint and a trap — it gives you the skeleton of the formula, but not exact figures. Hence the main challenge of manual analysis: a specialist sees the order but has to reconstruct plausible concentration ranges based on regulatory limits, the functional roles of ingredients, and typical formulation practice.

What Problems Does Competitor Composition Analysis Solve

Scientific diagram illustration showing a vertical bar chart or funnel visualization of cosmetic ingredients decreasing in concentration from top to bottom, labeled with generic terms like Aqua, Glycerin, Emulsifier, Preservative, with percentage scale on the side, clean infographic educational style, white background, blue and amber accent colors, no text overlap, no human figures
Comparing two products' formulas by phase and active ingredient

Studying someone else's formula is rarely driven by simple curiosity. In practice, it serves several concrete business and R&D goals:

  • Formula benchmarking — comparing your own recipe with the category's top products in terms of base type, actives, and supporting ingredients.
  • Spotting a product's weak points — identifying outdated preservatives, excessive silicone loading, or actives claimed on the packaging that aren't actually present at an effective concentration.
  • Reverse-engineering a reference texture — understanding which emulsifier system and rheology modifiers create the recognizable consistency of a bestseller.
  • Checking marketing claims — does the "with retinol" positioning match the ingredient's actual place in the list, or is it a token dose hovering around 0.01%.
  • Planning differentiation — finding an unoccupied formulation niche when an entire segment relies on the same set of three or four humectants.

These tasks used to be handled by hand: a chemist would open the label, check each ingredient against an INCI Decoder or CosIng reference, estimate functional groups, and log hypotheses in a spreadsheet. The process took anywhere from thirty minutes to several hours per product — and the quality of the result depended heavily on the individual specialist's experience.

The Role of AI Analysis in Breaking Down Formulas

AI analysis of ingredient lists doesn't replace chemical expertise, but it radically speeds up the initial data processing. A language model trained on regulatory databases and formulation terminology can, within seconds:

  1. Identify the functional role of each ingredient in the list (emulsifier, preservative, humectant, antioxidant, and so on);
  2. Group ingredients by emulsion phase — water, oil, and active;
  3. Estimate plausible concentration ranges based on list position and known regulatory limits (drawing on formulation research reproduced by Draelos, 2015, and subsequent reviews of formulating practice);
  4. Flag potentially problematic combinations — for example, the incompatibility of certain preservatives with a high pH.

This kind of AI analysis of a competitor's INCI list turns a static string of letters on the label into a structured hypothesis about the formula: an approximate architecture, phase distribution, and assumptions about the concentrations of key actives. A specialist then verifies and corrects that hypothesis — the AI produces a draft, not a final verdict.

⚠️ Important: the AI model works with probabilistic estimates, not exact lab data. Final concentration figures always need verification through stability testing, regulatory limits, and, where necessary, independent lab analysis.

In the following sections, we'll break down how the process itself works — from preparing the INCI list to crafting a prompt for the AI and interpreting the resulting report — and walk through the entire path step by step using a concrete example.

What INCI Is, and Why It's Not Just an Ingredient List

INCI (International Nomenclature of Cosmetic Ingredients) is a unified naming system for raw materials, adopted for labeling cosmetic products. Every name in this nomenclature is tied to a specific chemical structure or botanical source and is independent of the manufacturer's trade name. That means Glycerin on a French brand's packaging and Glycerin on a Korean brand's packaging are molecularly identical, regardless of what marketing materials call it. But that very standardization of names is what turns the ingredient list into a coded text: without understanding the ordering rules and threshold values, the list is easy to misread — sometimes exactly backwards.

The Descending-Concentration Principle

Scientific illustration comparing two cosmetic product formulations side by side as structured molecular diagrams, each split into color-coded phases (aqueous phase in light blue, oil phase in pale yellow, active ingredients in green dots), magnifying glass icon highlighting a key active ingredient concentration difference, clean flat vector infographic style, laboratory research aesthetic, no human faces
Descending order of ingredient concentration in an INCI list

The key INCI rule is that components are listed in descending order of concentration — but only up to a certain point. Under the European Cosmetic Regulation (EC) No 1223/2009, which governs cosmetics labeling across the EU, ingredients present at above 1% must appear strictly in descending order by mass fraction. That means the first 3–5 positions on the list are almost always the base of the formula: water, emollients, emulsifiers, humectants. If Aqua is listed first and the next position is Glycerin, the developer is already telling you the approximate ratio of the water and glycerin phases — without a single number on the label.

⚠️ Important: ingredient order is not a guarantee of exact proportions — it's only a ranking. The gap between position #2 and #3 could be 15% or 0.3% — the regulation doesn't require exact figures, only order.

The Hidden Threshold: What Happens Below 1%

Once a component's concentration drops below 1%, the strict-ranking rule stops applying. Regulation 1223/2009 allows the manufacturer to list such ingredients in any order — this covers preservatives, fragrances, colorants, and low-dose actives (peptides, extracts, vitamins). That's exactly why an expensive peptide complex often sits right next to an ordinary colorant like CI 19140 at the very end of the list — from position alone, there's no way to tell whether one is at 0.9% and the other at 0.01%.

This is critical when analyzing competitor formulas: if an AI tool or a human analyst treats the "bottom third" of the list as unimportant, it's easy to miss an active ingredient that the manufacturer deliberately placed below the 1% threshold to avoid disclosing its real concentration and role in the formula.

Masking Names: Parfum, Aqua, and Umbrella Terms

The INCI nomenclature allows so-called "umbrella" designations that hide dozens of chemical substances behind a single word:

  • Parfum / Fragrance — can include 50–200 individual aromatic components, some of which are potential allergens, required to be listed separately only above certain thresholds (in the EU, 26 named allergens must be individually labeled above 0.001% in leave-on products and 0.01% in rinse-off products).
  • Aqua / Water — formally a single molecule, but the actual water phase in premium products often includes thermal, micellar, or distilled water with a different mineral profile — none of which is reflected in the name.
  • Umbrella terms like Parfum let a fragrance formula be protected as a trade secret — the regulation permits withholding the exact composition of the aroma blend, citing protection of the perfumer's intellectual property.

This opacity isn't a loophole — it's a deliberate compromise between the consumer's right to information and protecting proprietary formulas. The regulatory framework is set out in detail in the text of Regulation (EC) No 1223/2009 of the European Parliament, and the systematization of nomenclature rules and their historical development are covered in cosmetic-regulatory literature (per Rathi, 2011; a similar breakdown of INCI nomenclature architecture appears in Bilal & Iqbal, 2020).

Why This Changes the Approach to Analysis

When AI parses a competitor's INCI list, it needs to account not just for word order but for the structural rules: where the "above 1%" zone ends, which terms hide complex blends, and which positions near the end of the list might actually be expensive actives rather than "trace residue." Without that context, the model will interpret the list linearly — as if position #15 were automatically less significant than position #5, which isn't always true for real-world formulas.

How AI 'Reads' the Chemical Structure of a Composition: The Mechanics of NLP Parsing for INCI

An INCI string on a label isn't coherent text with grammar — it's a flat list of comma-separated tokens where order encodes concentration, not narrative logic. For a language model, that's an unusual task: standard NLP pipelines are trained on natural speech with syntactic relationships, while a cream's ingredient list is more of a nomenclature catalog, structurally closer to a chemical database than to prose. That's why parsing INCI requires a specialized processing chain rather than a general-purpose chatbot.

Tokenization and Named Entity Recognition (NER)

The first stage is Named Entity Recognition (NER), adapted for the chemical domain. The model splits the string into candidate tokens and classifies each as a potential ingredient name, even when it spans multiple words (Sodium Ascorbyl Phosphate is one entity, not three). The difficulty is that INCI nomenclature uses Latin botanical names, trade abbreviations, numeric indices (PEG-100 Stearate), and hyphenated or bracketed compound constructions — standard tokenizers trained on news corpora often split such terms incorrectly.

To improve accuracy, domain-adapted models such as BioBERT or ChemBERTa are used, fine-tuned on chemical and biomedical texts — their embeddings capture chemical-nomenclature patterns better than general-purpose models (per Beltagy et al., 2019, in the context of SciBERT for scientific text). Regex rules are also applied to catch numeric suffixes (-100, -20) and CAS-like patterns, which reduces false positives on unknown tokens.

Matching Against Databases: From String to Structure

After tokenization, each recognized entity goes through entity linking — matching against reference databases. Several sources are used simultaneously here:

  • CosIng (Cosmetic Ingredient Database) — the EU's official database, containing the INCI name, CAS number, functional class, and regulatory restrictions.
  • PubChem — a chemical-compound database with molecular structures, letting the model link an INCI name to a specific SMILES notation and physicochemical properties.
  • Internal synonym dictionaries, since a single ingredient may appear under different trade names from different suppliers.

Technically, the matching is implemented through vector similarity of embeddings (cosine similarity) between the recognized token and database entries, rather than exact text matching — this matters because real labels contain typos, inconsistent casing, and spelling variations (Tocopherol vs Tocopheryl Acetate are different substances that are easy to confuse with a surface-level string comparison). This fuzzy-matching, embedding-based retrieval approach is described in entity-linking work for biomedical terms (Wang et al., 2021, in the context of linking medical entities to UMLS).

⚠️ Important: automatic matching breaks down on rare plant extracts and proprietary complexes (for example, "Symwhite 377" as a trade name instead of the INCI name Phenylethyl Resorcinol). Such cases require manual verification — the model doesn't replace checking against primary sources.

Clustering by Functional Class

The final step is assigning each recognized ingredient a functional tag: emollient, surfactant, preservative, chelator, antioxidant, and so on. This isn't just a lookup in the CosIng table (even though the database has a function field) — it's often a multi-class classification task, since a single ingredient can perform several roles at once. Glycerin is a humectant, a solvent, and, at high concentrations, a preservative booster; the model has to weigh context (list position, neighboring ingredients) to determine the dominant function in a specific formula.

Processing stageMethodData source
TokenizationRegex + domain rulesInternal INCI-pattern dictionary
NERBERT-like models (BioBERT, ChemBERTa)Fine-tuning on chemical corpora
Entity linkingCosine similarity of embeddingsCosIng, PubChem
ClusteringMulti-class classificationCosIng functional tags + position context

This multi-stage architecture explains why a good AI INCI breakdown doesn't just produce a list of "what is this substance," but a map of each component's functional role in the formula — that's precisely what turns a raw list into an analytical tool. For more on how these functional clusters turn into a finished formula, see our piece on reformulating cosmetics from a competitor's INCI list.

A Step-by-Step Algorithm: From Grabbing a Competitor's Composition to a Finished AI Report

Breaking down an INCI list comes down to five technical steps, each of which affects the quality of the final report. Skipping any step — say, working with raw, unprocessed text copied from the packaging — feeds the AI "dirty" data, and what comes out is a set of generic phrases instead of concrete analysis.

Step 1. Collecting the Source Composition Text

A competitor's composition can be pulled from three sources: a marketplace product listing, the brand's official website, or a photo of the label. Priority goes to the manufacturer's website — the text there is usually proofread and unaffected by copy-paste errors. Marketplace listings often contain seller typos (missing hyphens, merged words, Cyrillic characters swapped in for Latin ones), which matters a lot for parsing.

  • Brand website — copy the text directly from the HTML ingredient block, avoiding screenshots.
  • Marketplace — cross-check against packaging photos if the listed composition is under 15 ingredients (a sign of a truncated description).
  • Label photo — recognize it via OCR (Google Lens, Adobe Scan) and always proofread the result by hand: OCR confuses Cetearyl Alcohol with Cetyl Alcohol, "I" with "l," and "O" with "0."
⚠️ Important: never feed a composition to the AI without checking for typos first. One mixed-up letter turns an INCI name into a nonexistent substance, and the model will either invent properties for it or honestly say it can't find it — either way, the report suffers.

Step 2. Normalizing the Text

Before sending it to the model, the composition is brought into a consistent format:

  1. Separate ingredients with commas, remove semicolons and line breaks within a single name.
  2. Normalize casing throughout — not critical for meaning, but it lowers the risk of the model "not recognizing" a name because of all-caps.
  3. Strip out footnotes and marketing annotations like "*organic ingredient" that sometimes get inserted directly into the INCI list.
  4. Check the order — under the regulation, INCI runs in descending concentration down to 1%, after which the order is arbitrary. That matters for interpretation rather than normalization, but it's worth noting at this same step.

The result is a clean string like: Aqua, Glycerin, Niacinamide, Butylene Glycol, Sodium Hyaluronate, Panthenol, Phenoxyethanol, Ethylhexylglycerin.

Step 3. Writing the Prompt

The prompt determines the depth and structure of the report. A weak request ("tell me about this composition") produces a generic rundown of ingredient functions that you could find in any database. A working prompt sets the role, the output structure, and the analysis criteria.

A solid prompt structure is built from four blocks:

Prompt blockContent
Role"You are a cosmetic formulation chemist experienced in reverse-engineering ingredient lists"
Input dataNormalized INCI list + product category (serum, cream, emulsion)
TaskDetermine the formula type (emulsion/gel/anhydrous), the actives and their likely concentration, the preservative system, and weak points
Output formatA table of "ingredient — function — list position — estimated %," plus a separate block with conclusions on stability and cost

Example working prompt: "Analyze this INCI composition as a formulation chemist. The product is a hydrating serum. Determine: 1) the base system type, 2) the three key actives and their approximate concentration based on list position, 3) the preservative system composition, 4) potential stability or compatibility issues. Output it as a table, then add a 5–7 sentence written conclusion."

Step 4. Initial Interpretation of the Output

The AI produces hypotheses, not facts — keep that in mind while reading the report. A good sanity check for the output:

  • The sum of the estimated active percentages shouldn't exceed 100%, and should logically track the ingredient's position in the list.
  • The stated preservative concentration should fall within regulatory limits — for example, Phenoxyethanol is capped at 1% under EU rules.
  • If the model calls an ingredient near the end of the list (past position 15) a "key active at a high concentration," that's a misreading of INCI ordering, and the conclusion needs to be corrected by hand.

Step 5. Assembling the Report

The final report is assembled from three parts: the ingredient breakdown table (from step 3), the model's written conclusions (manually checked per step 4), and the chemist's own commentary — comparison with reference formulas, a cost estimate, assumptions about product positioning. That third part is what turns the AI's output from generic analysis into an actionable competitive report, ready to brief a formulation project.

Decoding Functional Roles: What's Behind Each Ingredient

Once the AI has recognized the individual INCI tokens, a harder task begins: figuring out why each component is in the formula in the first place. An ingredient's name alone doesn't tell you its function: Cetearyl Alcohol can act as either an emulsifier or a structurant, while Citric Acid can be a chelator, a pH adjuster, and a mild exfoliant all at once. The AI resolves this ambiguity by tying molecular structure to physicochemical parameters — polarity, HLB value (Hydrophilic-Lipophilic Balance), and the presence of functional groups capable of complexation or redox reactions.

From Structure to Function: Three Parameters the Model Calculates

For each recognized ingredient, the model draws on a database of molecular descriptors — atomic weights, hydrocarbon chain length, the number of hydroxyl and carboxyl groups, the presence of aromatic rings. This data lets it estimate the logarithm of the partition coefficient (logP), which correlates with HLB and predicts how the molecule behaves at the phase interface (following the approach described in Griffin, 1954, and further developed in QSAR modeling of cosmetic surfactants — Klein et al., 2019).

ParameterWhat it showsExample ingredientAI's functional conclusion
HLB 3-6lipophilic balanceGlyceryl StearateW/O emulsifier
HLB 10-18hydrophilic balancePolysorbate 20O/W emulsifier, solubilizer
logP < 0high polaritySodium Ascorbyl Phosphateantioxidant, water phase
presence of carboxyl groups + low molar massability to chelate metalsDisodium EDTAchelator, stabilizer

Emulsifiers and HLB Balance: How AI Works Out the System Type

When the model sees a cluster of several ingredients with different HLB values, it reconstructs not just a list but the logic of the emulsion system. For instance, the combination of Cetearyl Alcohol (HLB≈4, co-emulsifier) and Ceteareth-20 (HLB≈15.5) is typical of an O/W emulsion with a phase working temperature of 70-75°C — the AI recognizes this pattern as a "classic binary emulsifying system" and matches it against typical industry formulation templates (see also our breakdown of emulsifiers in the article on choosing emulsifiers). If the composition includes a hydrocolloid like Xanthan Gum alongside only a small share of an actual emulsifier, the model concludes that stabilization comes from increased viscosity rather than classic surfactant interaction — these are fundamentally different approaches to reverse-engineering texture.

Recognizing Chelators and Antioxidants by Functional Groups

Chelators are identified by the presence of multiple carboxyl or phosphonate groups capable of forming coordination bonds with Fe³⁺ and Cu²⁺ ions — precisely the metals that catalyze lipid oxidation and vitamin degradation in a formula (Fessi & Puisieux, 2001, in the context of emulsion stability). Antioxidants are recognized differently: the model looks for aromatic hydroxyl groups (phenols) or redox-active phosphate-ester structures of ascorbic acid. So Tocopherol and Sodium Ascorbyl Phosphate don't look alike by name, but they're structurally united by their ability to donate an electron, which the AI tags under a single label — "primary antioxidant" — regardless of solubility.

⚠️ Important: automatic function classification runs on a probabilistic model and can be wrong on rare derivatives and proprietary complexes — an INCI name like Aqua (and) Glycerin (and) Alcohol (and) Sodium Citrate (and) Xanthan Gum can mask a multi-component premix where the real function comes from synergy rather than a single substance. It's always worth checking the AI's final report against typical dosages: chelators rarely exceed 0.2%, antioxidants run 0.5-1%, and emulsifiers 2-8% depending on the system type.

Why This Matters for Formula Reverse-Engineering

Determining function rather than just identifying a name is what turns an INCI list into a working formula map. Understanding that a given ingredient is a chelator, not just a "texture additive," tells you why it's placed right after water and glycerin (INCI order correlates with mass concentration, except below 1%, where the order is arbitrary — a regulatory rule set out in Regulation (EC) No 1223/2009). That lets you reconstruct not just the composition but the process logic: the temperature windows for phase addition, the sequence for adding actives, the expected pH range of the finished system — everything you need for a thoughtful reproduction of a competitor's formula rather than a mechanical copy.

Common Mistakes in AI Analysis of a Competitor's Composition, and How to Avoid Them

AI produces a nicely formatted composition report quickly, and that very speed creates a false sense of reliability. In practice, even a good model regularly stumbles in three areas: misreading ingredient order, fabricating details for unknown names, and confusing INCI synonyms. Let's look at each trap separately, with concrete examples of where AI typically trips up.

Mistake 1: Reading Ingredient Order Literally as Exact Concentration

The regulation only requires descending order down to the 1% threshold; after that, the order can be arbitrary (usually alphabetical). AI, especially with an under-detailed prompt, tends to treat the entire string as a strict hierarchy from highest to lowest — so a preservative sitting between plant extracts on the list gets described, by default, as "present at a meaningful concentration," even though it's really a fraction of a percent.

⚠️ Important: if ingredients that are obviously "expensive" or active (peptides, retinoids, niacinamide) show up past roughly position 7-10, that's almost always a sign of a sub-threshold concentration (under 1%) — a marketing "alibi" on the packaging rather than a working dose. Explicitly tell the AI to account for this in the prompt, or the report will overestimate the real effectiveness of the competitor's formula.

Mistake 2: Hallucinations on Unknown or Rare INCI Names

When the model encounters a nonstandard, outdated, or regional INCI name (say, a local plant extract registered only under an Asian nomenclature), it sometimes doesn't say "I don't know" — instead it generates a plausible-sounding but invented description of the function and origin. That's a classic LLM hallucination: the model fills the gap with something statistically likely rather than a verified fact.

Sign of hallucinationHow to check
An overly generic description ("a plant-derived moisturizing ingredient")Cross-check manually against CosIng or the INCI Dictionary
No specific function given (antioxidant/emollient/surfactant)Ask the AI for the source of its classification and to flag uncertainty
The name can't be found in open databasesCheck the spelling — it's often a transliteration error or typo in the source composition

Working rule: any ingredient for which the model can't give a specific functional role and a typical dosage range needs manual verification — don't include it in the report's final conclusions as an established fact.

Mistake 3: Confusing INCI Synonyms and Trade Names

The same substance can appear under different INCI names depending on the raw-material supplier, while colloquial trade terms (like "hyaluronic" or "vitamin C") hide dozens of chemically distinct derivatives with different stability and bioavailability. AI sometimes confuses Sodium Hyaluronate with Hydrolyzed Hyaluronic Acid, or lumps every ascorbic-acid derivative under a single "vitamin C" label without distinguishing Ascorbic Acid, Sodium Ascorbyl Phosphate, and Ascorbyl Glucoside — even though they have fundamentally different storage stability and optimal pH.

  • Always ask the AI for the exact full INCI name, not a generic category ("this is a form of vitamin C, but which one specifically?")
  • Ask the model to explicitly spell out the differences between similar names when the composition includes several derivatives of the same active
  • Don't trust conclusions about two compositions being "identical" just because their general ingredient categories match — check the actual INCI strings

How to Minimize the Risk of Errors in Practice

Three measures bring the error rate down to an acceptable level:

  1. Double-checking — manually verify any conclusion about an "active ingredient at a working concentration" against open databases (CosIng, PubChem), especially if it drives a formulation decision
  2. Explicit prompt instructions — ask the model up front to flag uncertain answers with "insufficient data" instead of generating plausible-sounding text
  3. Cross-checking with a second query — ask the same question with different wording, or run it through a different model, and compare the answers; a discrepancy is a signal for manual review

AI analysis of INCI is an accelerator for initial review, not a substitute for chemical expertise. The AI's report should be treated as a draft set of hypotheses that need confirming before they go into developing your own formula.

From Analysis to Action: Using AI Insights in Your Own Formula Development

An AI report isn't a finished product — it's raw material for decisions. The difference between a hobbyist formulator and a professional is that the former copies a competitor's composition almost verbatim, while the latter extracts patterns from the analysis and translates them into their own formula logic — with different ratios, a different cost, and a different claim set. Let's break down exactly how to turn dry parsing data into working decisions.

Benchmarking: Finding the Gap Between Claim and Reality

The first thing structured AI-analysis data gives you is the ability to compare several competitor compositions side by side. If you've run 5–7 products from the same category through the AI (say, serums with Niacinamide), you end up with statistics on where the actives sit in the INCI list — which directly correlates with their concentration.

BrandNiacinamide's INCI positionEstimated concentrationStated claim
Competitor A4th8–10%"Evens out tone"
Competitor B9th2–3%"With niacinamide"
Competitor C3rd10–12%"Anti-pigmentation"

A table like this immediately reveals the gap: Brand B is selling marketing, not efficacy. That's your niche — if you launch with an honest 8–10% and a substantiated claim, you'll position yourself between A and C, but with a more attractive price point thanks to optimizing the rest of the formula.

Finding a Niche Through Functional-Layer Differentiation

The AI is good at organizing ingredients by functional role (emollients, gelling agents, preservatives, antioxidants). Once you see the role distribution across 5+ competitors, it becomes clear where the market is saturated and where there's a gap.

  • Oversaturation — if every competitor uses the same gelling agent (Carbomer) with a similar texture, differentiating on texture alone won't work — the market is already saturated there.
  • Gap in barrier support — if none of the competitors have ceramides or cholesterol in their top 10 ingredients, and your audience has sensitive skin, that's an entry point.
  • Outdated preservation — finding Phenoxyethanol + Methylparaben in most competitors opens up a niche for a formula with a more modern system like Sodium Benzoate + Potassium Sorbate, which doubles as a "paraben-free" marketing claim.

Optimizing Cost Without Losing Function

Reading an ingredient's INCI position alongside known raw-material market prices lets you estimate a competitor's approximate cost per batch. If the AI determines that an expensive peptide complex sits at position 12 (meaning its concentration is a fraction of a percent — an "alibi dose"), you can decide to skip it entirely or replace it with a cheaper analog with a similar mechanism of action — for example, using Sodium Ascorbyl Phosphate instead of unstable pure ascorbic acid, keeping the antioxidant claim while simplifying formula stabilization.

⚠️ Important: cutting costs on an active shouldn't turn into a false claim. If you state "with retinol," make sure the concentration in your formula is clinically meaningful (usually from 0.1%), not a homeopathic dose there just for the marketing line.

From Insight to Prototype: Three Steps for the Formulator

  1. Lock in a hypothesis. Based on the AI report, formulate one concrete difference: "my formula will contain 10% Niacinamide instead of the market average of 3%" or "I'll remove denatured alcohol from the base and replace it with a hydrating complex."
  2. Recalculate the phase balance. Swapping or increasing active concentrations almost always requires revisiting the emulsifier and thickener — use the report as a compatibility checklist, not a ready-made recipe.
  3. Test on a small batch. A 100–200 g test batch with pH control (the target range for most serums is 5.0–6.0) and visual stability checked at 48 hours at room temperature and 45°C is the minimum bar before scaling up.

This approach turns competitor AI analysis from reconnaissance into a design tool: you're not reproducing someone else's formula, you're building your own with a justified set of actives, a realistic cost, and a claim that will hold up — both on the ingredient side and legally. For more on calculating batch cost and picking a preservation system for a specific formula type, see our piece on calculating the cost of a cosmetic formula.

Conclusion: AI as a Competitive-Intelligence Tool, Not a Replacement for a Formulation Chemist

Breaking down a competitor's INCI composition with AI is work with probabilities, not facts. The language model offers the most plausible interpretation of ingredient order, concentrations, and functional roles, drawing on statistical patterns in its training data. It never opened the packaging, never ran a titration, and never saw the brand's actual formula. Everything it produces is a hypothesis that needs to be checked against chemical knowledge — not a final verdict.

That's exactly why the "AI + formulation chemist" combination works better than either tool alone. The AI takes over the routine work: sorting components by the descending-concentration rule, matching INCI names against functional-class databases, and generating a draft hypothesis about emulsion type and the formula's approximate price segment. The chemist takes over what the model fundamentally can't do — checking chemical compatibility, assessing stability at a given pH and temperature, and calculating realistic dosages within regulatory limits.

What AI Does Well, and What It Doesn't

TaskAI's roleChemist's role
Decoding an ingredient's function from INCIFast initial classificationChecking context (concentration, neighboring ingredients)
Estimating a concentration rangeA rough bracket based on list positionRefinement based on regulatory limits and hands-on raw-material experience
Hypothesizing the emulsifying system typePicking the most likely surfactant pairChecking compatibility with the formula's pH and electrolytes
Assessing finished-formula stabilityNot performed reliablyCalculation and lab testing are mandatory

This division of labor isn't a compromise — it follows from the nature of the tool. The AI is trained on textual patterns, not the physical chemistry of emulsions. It can "see" Cetearyl Alcohol and Polysorbate 60 in a list and recognize a typical emulsifying pair, guessing an oil-in-water system, but it can't verify whether that system stays stable once you add 2% Niacinamide at pH 5.0. Only a specialist with an understanding of phase diagrams and hands-on stability-testing experience can do that check.

⚠️ Important: an AI report on a competitor's INCI composition is a draft for discussion with a chemist, not a production spec. Launching a batch based on an unverified AI hypothesis about concentrations and ingredient compatibility risks an unstable formula, consumer allergic reactions, and regulatory action.

Practical Takeaway

Framed as a working protocol, the takeaway looks like this:

  1. Use AI for a fast initial decoding of a competitor's composition — it saves hours of manual work with INCI references.
  2. Treat any concentration figure from the AI report as a range, not an exact value, and cross-check it against regulatory limits (INCI Dictionary, CIR reports, EU Regulation 1223/2009).
  3. Verify hypotheses about emulsion-system type and ingredient compatibility with HLB calculations, stability testing, and physical-chemistry knowledge — this step can't be delegated to the model.
  4. Use the analysis's insights not to copy a competitor's composition, but to build your own differentiated formula with a justified set of actives.

Competitive intelligence through AI INCI analysis is an accelerator at the start of development, cutting preliminary market analysis from days down to hours. But the final call on active-ingredient concentration, preservative choice, or emulsifying system stays with the formulation chemist, who bears responsibility for the finished formula's stability, safety, and regulatory compliance. AI extends a specialist's reach, but it doesn't replace fundamental knowledge of cosmetic chemistry — and that balance is what determines whether the AI tool becomes a useful assistant or a source of costly development mistakes.

Breaking down other people's formulas is useful, but building your own is even more useful. AI Chemist in the Walker Formulation Academy Club helps with both: it's tied to a verified INCI database, so it doesn't confuse ingredient functions and working ranges the way a generic chatbot does. And in live sessions with a real instructor, you'll learn to read compositions like a professional.

Walker Formulation Academy Club

Enjoyed the article? Get access to the AI Chemist and video recipes

The 24/7 AI assistant answers formulation questions, calculates HLB and pH and helps you choose ingredients. Plus a private community of chemists and monthly product reviews.

No card required · Cancel anytime

Rate this article

Your rating helps other readers find useful guides

How to Break Down a Competitor's INCI List with AI: A Step-by-Step Guide | Walker Formulation Academy