How do you know if a laundry detergent really performs better?
Two laundry detergents can produce different cleaning results in a test, but that does not automatically mean one product genuinely performs better than the other.
Meaningful detergent performance testing requires more than washing a few stained fabrics and comparing the numbers. The stains selected, wash conditions, number of replicates, measurement method and statistical analysis can all influence the conclusion.
So how do you determine whether an apparent performance advantage is real?

A higher result does not necessarily mean better performance
When detergent performance is measured, some variation between individual test results is inevitable. Fabric, stain application, washing conditions and measurement itself can all introduce normal experimental variation.
For example, if Product A produces an average soil-removal result of 62% and Product B produces 60%, it may be tempting to conclude that Product A is the better detergent.
But a two-percentage-point difference means very little without understanding the variability behind those averages.
If the results vary substantially between replicates, the apparent difference may simply reflect normal test variation. If the results are highly repeatable, however, even a relatively small difference may be meaningful.
This is why comparative detergent testing should consider both the size of the observed difference and the confidence that the difference is genuine.
Performance testing should answer whether products are meaningfully different — not simply identify which average is numerically highest.
Start with the right stains and soils
The stains selected for a laundry detergent test have a major influence on what the test will actually tell you.
A detergent may perform very well on oily soils but less effectively on protein-based stains. Another product may show the opposite pattern. If the stain set is too narrow, the resulting comparison can therefore give a misleading impression of overall performance.
A useful laundry performance test should include a representative range of stains that reflect the types of cleaning challenges relevant to the product and market being evaluated.
Typical stain groups might include:
- Oily and greasy soils
- Protein-based stains
- Starch and carbohydrate soils
- Particulate soils
- Pigmented or coloured stains
- Everyday food and beverage stains
The objective is not necessarily to use the largest possible stain set. It is to choose a set that provides enough diversity to reveal meaningful differences between the products being tested.
For product-development work, stain selection may also be targeted. If a formulation is known to struggle with oily soils, for example, the test programme can include more challenging grease-related stains so formulation changes can be evaluated specifically against that weakness.
The stain set defines the performance question. If the wrong stains are selected, even a perfectly executed test may answer the wrong question.
Control the wash conditions
Even with the right stain set, a detergent comparison can become misleading if the wash conditions are not tightly controlled.
Laundry performance is influenced by factors such as water hardness, wash temperature, detergent dosage, machine type, programme, load size and wash time. If these conditions vary between products, it becomes difficult to know whether the observed difference came from the formulation or from the test itself.
A meaningful comparison therefore requires the relevant variables to be defined in advance and applied consistently across all products.
Water hardness
Detergent performance can change significantly with water hardness. Builders, surfactants and other formulation components may respond differently as calcium and magnesium levels increase, so the selected water hardness should reflect the market or use condition being investigated.
Temperature
Wash temperature affects soil removal, surfactant behaviour, enzyme activity and other aspects of detergent performance. Products should therefore be compared at the same controlled temperature unless temperature response itself is the subject of the study.
Dosage
Using inconsistent detergent doses can overwhelm genuine formulation differences. Dosage should be defined clearly and applied accurately across the comparison.
Wash programme and equipment
Machine type, cycle duration, agitation, water volume and other programme characteristics should remain consistent so each detergent experiences the same mechanical and process conditions.
The purpose of controlling the wash conditions is simple: when the results differ, you want the detergent to be the reason.
Replication matters
A single wash result can be useful during early development, but it is rarely enough to conclude that one detergent genuinely performs better than another.
Even under carefully controlled conditions, some variation between individual swatches and wash runs is unavoidable. Differences in stain application, fabric, machine behaviour and measurement can all influence the final result.
Replication helps separate these normal sources of variation from real differences between the detergents being compared.
More measurements give greater confidence
By testing multiple stained swatches and, where appropriate, repeating the wash comparison, it becomes possible to estimate how variable the results are for each product.
A detergent that consistently performs better across repeated measurements provides much stronger evidence than one that happens to produce the highest result in a single test.
Replication also becomes particularly important when the difference between products is relatively small. A large performance difference may be obvious, but smaller differences require enough data to determine whether they are repeatable.
Replication should suit the question
There is no single number of replicates that is correct for every detergent study. The appropriate level depends on factors such as the variability of the method, the size of the expected difference and how the results will be used.
Development screening may require a different level of replication from competitor benchmarking or testing intended to support a comparative performance claim.
One result tells you what happened in that test. Replication helps tell you whether it is likely to happen again.
Use statistics to determine whether the difference is meaningful
Once you have controlled the wash conditions and generated enough repeated measurements, the next question is whether the observed difference between products is statistically meaningful.
Averages alone do not tell the whole story.
For example, Product A may achieve an average soil-removal result of 68% while Product B achieves 65%. That three-point difference could represent a real performance advantage — or it could simply sit within the normal variability of the test.
Statistical analysis helps determine how confidently the observed difference can be interpreted.
Statistical significance
A statistically significant result suggests that the observed difference is unlikely to be explained by normal test variation alone.
This is particularly useful when comparing products that perform relatively closely, where visual inspection of the average results may not provide a reliable conclusion.
Equivalence matters too
Not every comparison needs to identify a winner.
In many development or benchmarking studies, it can be equally important to establish that two products perform similarly within the limits of the test.
That may be relevant when:
- reformulating a product to reduce cost;
- replacing a raw material;
- developing an alternative formulation;
- assessing whether a prototype has reached benchmark performance;
- or comparing several products within a market category.
Statistical significance is not the same as commercial significance
A difference may be statistically significant without being large enough to matter commercially.
Conversely, a commercially important difference may require more data before it can be demonstrated with statistical confidence.
This is why detergent performance results should be interpreted in the context of both the statistical evidence and the practical objective of the study.
Statistics help determine whether the difference is likely to be real. Technical interpretation determines whether that difference actually matters.
Look beyond the overall average
An overall performance score can be useful for summarising a detergent comparison, but it can also hide important differences between products.
Two detergents may achieve almost identical average cleaning scores while performing very differently across individual stain types.
For example, one product may be particularly strong on oily and greasy soils but weaker on protein-based stains. Another may perform more consistently across the whole stain set without being the strongest on any individual stain.
If only the overall average is considered, those differences can disappear.
Performance profiles can reveal more than rankings
Looking at results stain by stain helps build a performance profile for each detergent.
This can show:
- where a product has a genuine performance advantage;
- where competitors are stronger;
- whether weaknesses are concentrated within particular soil types;
- whether a reformulation has improved one area while reducing performance elsewhere;
- and where future formulation work is likely to deliver the greatest benefit.
This is especially valuable during product development. A formulation does not necessarily need every stain result to increase. The objective may instead be to address a known weakness while protecting areas where the product already performs strongly.
Category averages can also be useful
Where a test contains a broad range of stains, results can sometimes be grouped into meaningful categories such as oily soils, protein stains or particulate soils.
This can make patterns easier to interpret without reducing the entire test to a single number.
A detergent’s average score tells part of the story. Its performance profile tells you where the strengths and weaknesses actually are.
Interpret the results in the context of the question
The same set of detergent performance results can mean very different things depending on why the testing was carried out.
A product-development study, for example, may be trying to identify where a prototype still needs improvement. A competitor benchmark may be asking whether a product is genuinely stronger than the market leader. A reformulation project may simply need to demonstrate that performance has been maintained after a raw-material or cost change.
This is why the technical question should be defined before the test programme is designed.
Product development
During formulation development, the most useful result is often not an overall ranking but an understanding of where the formulation is improving and where weaknesses remain.
Performance data can help determine which formulation changes should be retained, which should be reversed and where the next round of development effort should be focused.
Competitor benchmarking
When comparing commercial products, the objective may be to understand where a product sits within the market.
That means identifying not only where one product performs significantly better, but also where products are statistically similar and which stain types are responsible for the most important differences.
Reformulation and cost optimisation
If an existing detergent is being reformulated to reduce cost, replace a raw material or meet a new ingredient requirement, the question may be whether the revised formulation has maintained acceptable performance rather than whether it has become the highest-performing product in the category.
Claims support
Where testing is intended to support a comparative performance claim, the test design, benchmark selection, replication and analysis need to be appropriate to the specific statement being considered.
A result that is useful for internal development does not automatically provide sufficient evidence for a marketing claim.
Good performance testing begins with a clear question and ends with a conclusion that answers that question.
So, which detergent really performs better?
The answer cannot reliably be determined from a single wash, one overall average or the highest number on a chart.
A meaningful laundry detergent comparison requires the products to be tested under controlled and relevant conditions, across an appropriate range of stains, with enough replication to understand normal variation. Statistical analysis can then help determine whether apparent differences are likely to be genuine.
But even that is only part of the answer.
The results also need to be interpreted in the context of the question being asked. A detergent may be stronger overall, perform particularly well on certain soil types, be statistically equivalent to a benchmark, or deliver an improvement that is technically real but commercially insignificant.
Good performance testing therefore brings together:
- Relevant stains and soils
- Controlled wash conditions
- Appropriate replication
- Objective measurement
- Statistical analysis
- Technical interpretation
The result should be more than a ranking. It should provide a clear understanding of where products differ, whether those differences are meaningful and what the findings mean for the next decision.
Knowing which detergent produced the highest average is easy. Knowing whether it genuinely performs better requires a properly designed experiment.
Need to compare laundry detergent performance?
d-labs designs and conducts independent laundry detergent performance testing for product development, competitor benchmarking, market evaluation and claims support.
