Cosmetic stability testing is the set of laboratory procedures that verify a cosmetic formula stays safe, effective, and visually unchanged over its declared shelf life, under realistic storage and use conditions.
It is the part of product development that brand founders understand the least and skip the most often.
It is also the part that causes the most expensive post-launch problems when it goes wrong.
A formula that smells different after three months on a shelf. A cream that separates in hot weather. A preservative that stops working after the customer opens the jar the tenth time.
All of these are stability failures, and all of them are preventable.
Most stability testing content online is written for lab technicians and formulators.
It goes deep into test protocols, temperature conditions, and measurement methods that the brand founder does not need to understand at that level.
After 30 years in the hair and beauty sector, most recently in private label cosmetics, I can tell you what you actually need to know to make decisions without becoming a lab expert.
This guide walks through stability testing from the brand founder’s perspective. Enough detail to make decisions, not so much that you get lost in the chemistry.
What Is Stability Testing and Why Does It Matter?
Stability testing answers three questions about a finished cosmetic product.
Does the formula stay the same over time? Does it stay safe once the customer opens the package? Is the packaging still compatible with what is inside it?
If the answer to any of those three is no, you have a product problem. Not a small one.
Stability is the bridge between the perfect sample you approved in the lab and the product that arrives in a customer’s home six months later.
A formula that performs beautifully on day one can fail on day 60 if it was not tested properly.
Those failures usually happen in three ways.
Chemical stability. Active ingredients degrade, losing efficacy. Vitamin C oxidizes. Retinol breaks down. The product you sell at month 12 is not the product the customer expected when they bought it.
Physical stability. The formula separates, settles, changes texture, or changes color. Emulsions break, creams turn grainy, and liquid serums go cloudy.
Microbiological stability. The preservative system fails, and the product grows bacteria, yeast, or mold. This is the scenario that causes product recalls, consumer skin reactions, and regulatory action.
Stability is the difference between a product that builds a brand and a product that destroys one.
For how stability testing fits into the full regulatory picture you build alongside the product, see the Product Information File guide.
The Three Types of Stability Testing You Need to Know
There are three main test families. Each answers a different question. Most products need all three.
Real-time stability testing
Real-time stability (also called long-term stability) is the simplest of the three.
The product is stored at normal conditions (25 degrees Celsius, 60% relative humidity) for the full declared shelf life, usually 24 to 36 months.
At set intervals (3, 6, 12, 18, 24 months), samples are evaluated for changes in color, smell, texture, pH, viscosity, and chemical composition.
This is the most reliable form of stability testing because it simulates exactly what will happen to your product on a shelf.
The problem is obvious: waiting 24 months before launching a product is not commercially viable for most brands.
Accelerated stability testing
Accelerated testing is the workaround for the real-time problem. The product is stored at elevated temperatures (typically 40 or 45 degrees Celsius, sometimes 50) for 3 to 6 months.
The principle: higher temperatures speed up chemical and physical changes.
Three months at 40 degrees is roughly equivalent to 12 to 18 months at room temperature for many formulas.
This is what a manufacturer uses to justify a 24-month shelf life after 3 months of testing, with the real-time data collected after launch to confirm it, instead of waiting 2 years before selling anything.
The trade-off is accuracy. Accelerated testing is a prediction, not a certainty.
Certain formula types (natural-heavy, biotech-heavy, some fragrance-driven products) do not age linearly, and accelerated results can over- or under-predict real shelf life.
Good manufacturers know which formulas need supplementary real-time data to be reliable.
Compatibility testing (product + packaging)
Also called package compatibility testing or PCT. This tests whether the formula and the chosen packaging interact badly over time.
Does the formula leach colorants from a plastic bottle? Does it react with a metal pump spring?
Do packaging plasticizers migrate into the product, and does the packaging material change its smell?
Compatibility testing is specific to the exact packaging you are using.
Change the bottle supplier, change the pump, change the cap material, and the test needs to be rerun.
It is one of the most commonly skipped tests, and one of the biggest sources of post-launch problems.
A formula that performs perfectly in a glass jar can fail within weeks inside a specific plastic tube.
Other testing commonly discussed alongside stability
Two more tests come up in almost every stability conversation, even though they are technically separate from stability testing itself.
Preservative efficacy testing (PET, also called challenge testing). This one measures contamination resistance rather than change over time. The product is deliberately inoculated with known microorganisms and monitored to see whether the preservative system kills them within the required timeframe. Annex I, Part A, point 3 names the preservation challenge test among the microbiological quality data, and in practice it is waived only where a low microbiological risk is justified.
Period-After-Opening (PAO) determination. This decides the "6M" or "12M" symbol, the one showing how long the product stays safe after the customer opens it. Article 19(1)(c) calls for that indication on products with a minimum durability of more than 30 months, unless durability after opening is not a relevant concept. The period itself depends on preservative efficacy plus use-pattern testing.
Cost, Timeline, and Who Is Responsible
This is where most founder confusion lives, and where most preventable mistakes happen.
The brand founder’s first question is usually: what does stability testing cost me? The honest answer depends on two things: which production path you are on, and who the contract assigns the work to.
Who pays depends on which production path you are on.
In the white label path, the manufacturer’s existing catalog formula already has stability data. You pay nothing extra because the testing is already done.
In the hybrid approach, the base formula was tested, but your customization may or may not be covered depending on how much you modified.
Sometimes the existing data is sufficient. Sometimes a supplementary test is needed.
In full custom formulation, stability testing is always part of the project and appears as a separate line item on your quote.
The typical cost and timeline for each test, at the brand-founder level. All figures below are indicative estimates that vary by manufacturer, region, and project scope.
Test type
Typical cost (per product)
Timeline
Legal status
Real-time stability
800-2,500 EUR/USD
24-36 months
No test named by law; the data is required
Accelerated stability
500-1,500 euros
3-6 months
No test named by law; the data is required
Compatibility testing
300-1,200 euros
3-6 months
No test named by law; packaging data is required
Preservative efficacy (PET)
300-600 euros
4-6 weeks
Named by Annex I; waived only where low microbiological risk is justified
PAO determination
Usually bundled with PET
Same as PET
Only if durability exceeds 30 months
A typical first product launch budgets 1,500 to 5,000 euros total for stability-related testing, depending on how much is already covered by the manufacturer’s existing data.
For multi-SKU launches (3 to 5 products), the testing cost rarely scales linearly.
Shared test panels and combined protocols bring the per-product cost down, usually to 60-70% of the standalone price.
The cost of accelerated stability testing for a single product is 500 to 2,000 euros. The cost of a post-launch stability failure starts at 15,000 euros and has no upper limit.
Responsibility depends on the contract, not on common sense.
In most private label contracts, stability testing is the manufacturer’s responsibility to run but the brand’s responsibility to pay for and approve. The results are shared with you and go into your Product Information File.
If the contract is vague about this, you want to clarify before signing. Who runs the tests, who pays, who owns the documentation, what happens if the tests fail.
Can You Skip Stability Testing? When the Answer Is Almost Never Yes
Short answer: no. The data is not optional, and skipping it is not practical either.
In the EU, Annex I, Part A, point 2 of Regulation 1223/2009 requires the safety report to cover the stability of the product under reasonably foreseeable storage conditions. The Regulation asks for the result, never for a method, a temperature or a duration.
But without that information, the safety report inside your Product Information File is incomplete, and the product cannot legally go on the market.
The PIF is not a signed document. The only signature the Regulation expressly requires is the safety assessor’s, dated, on Part B of the safety report.
A file without stability data puts the Responsible Person in breach of Articles 10 and 11. Liability for consumer harm on top of that runs under national law, not under Regulation 1223/2009.
In the US, MoCRA requires safety substantiation, and stability is part of what substantiation means. The FDA can request stability data if they investigate a product, and not having it is not a defensible position.
Beyond regulation, skipping stability testing is a commercial decision that looks something like this.
Year one: you save 2,000 euros on testing, launch three months earlier, celebrate the lower cost.
Month six after launch: first customer complaint about smell change.
Month eight: customer complaint about separation.
Month ten: first return batch from a retailer because stock is showing discoloration.
Month twelve: you are recalling product, reformulating from scratch, losing shelf space, and paying damages.
I have seen brands try this. Every single one of them paid for the mistake in the next 18 months.
The brands that take stability testing seriously rarely talk about it. They built it into the timeline and the budget from the start. Skipping it is what turns a brand into a cautionary tale.
There is one semi-legitimate scenario where you can move faster.
If you are launching with the white label or hybrid approach and the base formula has existing stability documentation that covers your version, you can rely on that.
Your manufacturer will tell you clearly whether the existing data is valid for your specific product. If they say yes, document it in your PIF. If they say no, run the test.
Common Stability Testing Mistakes
A few traps I see repeatedly, worth flagging before you start.
Mistake 1: Treating stability testing as a last-minute task.
Stability testing takes months, not weeks.
A founder who decides to "sort out testing" after the formula is approved and the packaging is ordered will find their launch delayed by 3 to 6 months.
Fix: include stability testing in the timeline from the first call with the manufacturer. Know the timeline before committing to a launch date.
Mistake 2: Running stability on the wrong package.
A brand approves samples in a temporary lab bottle, runs stability on that bottle, and then launches in a different commercial bottle.
The compatibility data does not transfer.
If the new packaging reacts badly with the formula, you find out from customers, not from the lab.
Fix: run stability and compatibility testing on the final commercial packaging, not on lab samples. If the packaging is not finalized, you are not ready to start stability.
Mistake 3: Ignoring accelerated test failures because "it is only accelerated."
Accelerated testing is not perfect, but a clear failure in accelerated conditions is a strong warning signal, not a false alarm.
A formula that separates at 40 degrees often separates at room temperature in a summer shipping container.
Fix: treat accelerated failures as a signal that something in the formula or packaging needs to change before launch.
Mistake 4: Not retesting after any formula change.
Any ingredient change, concentration change, fragrance change, or preservative change invalidates previous stability data. The product needs to be retested, even if the change seems small.
Fix: any change after stability testing means the testing clock restarts. This is a good reason to finalize the formula fully before testing begins.
Mistake 5: Assuming "natural" or "clean" formulas need less testing.
The opposite is true. Natural and clean formulas often use alternative preservatives that are less reliable than traditional ones.
Natural botanicals can vary batch-to-batch and introduce stability variability.
These formulas usually need more testing, not less.
Fix: if your brand is natural or clean, raise the testing budget rather than trimming it, and make sure your manufacturer has experience with the specific preservative system you are using. For how formulation philosophy affects what testing you need, see the formulation types guide.
The takeaway
Stability testing never shows up in your brand story or in your product descriptions.
It still decides whether you sell the same formula for three years or recall it after six months.
Brands that put stability in the budget from the start rarely have a stability crisis. Where it is treated as an optional cost, the bill arrives later and larger.
Put it in the budget and in the timeline, ask your manufacturer what data already exists and what still has to be run, then keep the results in your PIF before you launch.
The only real shortcut here is the manufacturer’s existing data, and only where it genuinely covers your product.
Frequently Asked Questions
Is stability testing legally required for cosmetics?
The data is, in the EU, the UK, and in most regulated markets. Under EU Regulation 1223/2009, Annex I, Part A, point 2, the safety report has to cover the stability of the product under reasonably foreseeable storage conditions, and that report sits inside the Product Information File (PIF). The Regulation asks for the result, never for a method, a temperature or a duration, so no named test is itself a legal requirement. But without stability information the PIF is incomplete and the product cannot be legally sold. In the US under MoCRA, safety substantiation is required and stability is a core part of that. Skipping stability data is not a legal option for any serious market.
How long does stability testing take before I can launch?
At minimum, 3 months of accelerated testing plus 4 to 6 weeks of preservative efficacy testing, in parallel where possible. For a typical first launch, plan 3 to 4 months of testing time once the final formula and final packaging are locked. Add more if you need compatibility testing on complex packaging. Real-time testing continues for 24 to 36 months after launch, collected as your product is already on the market.
Can I use stability data from the manufacturer’s existing formula?
Sometimes yes, sometimes no. In white label, the existing data usually covers your product directly. In the hybrid approach, it depends on how much you modified the formula and whether the modifications affect stability-relevant parameters (pH, preservation, fragrance load, color). In full custom, you need new testing. Your manufacturer should tell you clearly which scenario applies and document it in your PIF.
What happens if my product fails stability testing?
You do not launch with that formula. The manufacturer analyzes what failed (ingredient degradation, formula separation, packaging interaction, preservative weakness) and reformulates to address the issue. This is normal, not catastrophic. A formula that fails accelerated testing has just saved you from launching a product that would have failed in the customer’s hands. The cost of reformulating before launch is a fraction of the cost of recalling afterward.
Do I need to test every product in my line separately?
Generally yes, but protocols can be combined. Each distinct formula needs its own stability data. Each distinct packaging needs its own compatibility data. However, a lab can run multiple tests in parallel, and costs per product drop when you test 3 to 5 SKUs together. Products that share a formula base with only fragrance variations may share some data, but verify this with your manufacturer case by case.
Who keeps the stability test documentation?
The Responsible Person keeps it, and it lives in the Product Information File. In most private label arrangements, the manufacturer runs the tests and provides the documentation, but Article 11 puts the file, and the duty to keep it, on the Responsible Person. Make sure you have copies of all stability data before your product launches, not after.
What is the difference between stability testing and challenge testing?
Stability testing verifies that the formula does not change unacceptably over time (chemical, physical, microbiological changes across months). Challenge testing, also called preservative efficacy testing (PET), verifies that the preservative system can kill contamination when introduced. Stability answers "does this product stay safe and effective over its shelf life?" Challenge testing answers "can this product handle being contaminated and still be safe?" For a water-based cosmetic the safety report needs both sets of data.
Honest assessment of AI in cosmetics from an industry veteran with 30 years of experience. What AI can do, what it cannot, and how to use it smartly in your brand.
How to request, evaluate, and give feedback on cosmetic product samples. A practical guide to the sampling process that prevents expensive production mistakes.
How to create a haircare product line that actually sells. From hairdresser-turned-consultant: product selection, professional vs retail, and launch strategy.