We took twenty marketing claims routinely made for cosmetic ingredients and graded each one against the human evidence for that specific claim — not for the ingredient in general.
Seven claims had no human data bearing on them. One was supported by a direct head-to-head trial. The remaining twelve rested on genuine human studies that measured something other than what the claim promises.
That last group is the interesting one, and it is the majority. The common assumption is that cosmetic marketing invents things. Mostly it does not. It takes a real measurement and describes it as a different one.
We are calling this outcome substitution, and the pattern is consistent enough to be useful as a reading habit.
Ingredient grading is not new. Several databases rate cosmetic ingredients for safety, and a few rate them for efficacy. What none of them does is grade the promise.
This matters because the promise and the ingredient are separate claims. Niacinamide is one of the best-evidenced cosmetic actives in existence. That says nothing about whether a specific product's claim about niacinamide is supported. The ingredient can be excellent and the sentence on the box can still be unearned.
So we graded the sentences.
Each claim was written in the form the industry uses it — not our paraphrase, not the scientific outcome. "Collagen-boosting." "Natural retinol." "As effective as retinol." Then each was graded against the human evidence bearing on that exact assertion.
The grading uses one axis, with quality living inside the lower band:
Grade D — human data exists, but the claim is not supported by it. Three ways this happens:
Grade F — no human data bears on this claim at all. Also three ways:
Grade C or above — the claim is directly supported, with the usual caveats about study size and independence.
One boundary is worth stating explicitly, because we got it wrong first: manufacturer-run human studies are D, not F. A small unreplicated trial by the ingredient supplier is weak evidence, but it is human evidence. Folding it into F would have collapsed the distinction the scale exists to draw.
A second boundary: for a comparative claim, the comparison is the promise. "As effective as X" without a head-to-head trial is not weakly supported — it is unsupported, because the specific assertion has never been tested. Those claims go to F regardless of how much evidence exists for the ingredient in isolation.
That rule is about the absence of a comparison, not its weakness, and two rows in this set show the difference. A concentration claim — that more of a well-evidenced active outperforms less — stayed at D because the comparison was actually run in humans and the higher concentration did not win. A peptide claiming to be roughly a third more active than its predecessor also stayed at D: the comparison exists, it was run by the manufacturer, and manufacturer human data is weak evidence rather than no evidence. Both are worse claims than a clean F would suggest in one sense — someone checked, and the answer was not what the box says — but they belong in a different category from claims nobody has ever tested.
Twenty claims across sixteen ingredients.
| Grade | Count | What it means |
|---|---|---|
| C | 1 | Direct head-to-head evidence exists |
| D | 12 | Human data exists; claim not supported by it |
| F | 7 | No human data bears on the claim |
Within the twelve D grades:
| Reason | Count |
|---|---|
| Measured a different outcome | 8 |
| Result falls short of the promise | 3 |
| Only poor-quality human evidence | 1 |
Eight of twenty claims follow the same shape. A real human study measures one thing. The marketing describes a different thing. Both statements are true in isolation.
The clearest case is collagen.
Several peptides are sold on the promise that they boost collagen synthesis. The human trials behind them measured the appearance of skin — fine lines, wrinkle depth, roughness, firmness as assessed by graders or instruments. Those trials often found real improvements. What they did not measure is collagen.
The gap is not rhetorical. Skin can look better for reasons that have nothing to do with collagen synthesis: hydration, barrier repair, light-scattering changes in the stratum corneum, reduced inflammation. A trial that measures appearance cannot distinguish between them. So "improves the appearance of fine lines" is earned, and "boosts collagen" is a mechanism asserted on top of an outcome that does not test it.
The same shape appears elsewhere:
In every case the study is real, the improvement is real, and the sentence on the box describes a different finding.
A smaller group has the right outcome measured and simply does not reach the promised size. A barrier ingredient with human data showing protection during use, marketed as lasting restoration. A peptide marketed on a percentage improvement larger than its own trial produced. A concentration claim — that more of a well-evidenced active is better — where the comparison was actually run and did not show the advantage.
That last one is worth pausing on. It is the only case in the set where the industry's comparison exists and came out against the marketing. Ten percent did not beat five. That is stronger than "nobody checked."
The seven F grades are the ones that would make the best headline, and they are the least instructive.
Two are simply factually wrong: a plant compound sold as "natural retinol" when it is not a retinoid and shares no structural relationship with one, and a claim of regulatory approval for a category that regulator does not approve.
Five are comparative claims where the comparison was never run — "as good as" and "equivalent to" assertions with no head-to-head trial behind them.
These are easy to spot and easy to condemn. But if you learn only from them, you learn the wrong lesson: that cosmetic marketing lies. Mostly it does not lie. It substitutes.
One claim in twenty earned a C: a plant compound marketed as being as effective as retinol, where a genuine head-to-head trial exists.
The trial is small — forty-four participants — and one trial at that size cannot settle equivalence. A difference not detected is not equivalence demonstrated. But the comparison was actually conducted, and that alone separates this claim from the five comparative F grades.
It is worth naming as the standard: this is what a supported comparative claim looks like, and how rare it is.
Twenty claims is a small set. They were selected as the claims most commonly attached to ingredients already in our register, which is not a random sample of the market. A larger and differently-selected set could shift the proportions.
The grades are our judgement. The criteria are published and the reasoning for each row is written out, so disagreement is possible and checkable — but two careful readers could grade some rows differently, particularly at the D/F boundary.
We did not survey the market. We graded claims as commonly formulated, not as printed on specific products. A given brand may be more careful, or less.
We sell cosmetic serums. Some ingredients graded here compete with our own products. Our own copper-peptide serum is graded no more kindly than the peptides it competes with — that is checkable in the same dataset. We state the interest so it can be weighed.
The practical version is a single question, and it works on any product:
What did the study measure, and is that the same thing the box promises?
If the box says a mechanism — synthesis, repair, regeneration, penetration — and the evidence is an appearance study, those are different claims. The appearance result may be perfectly real. The mechanism is an addition.
This is not a reason to distrust cosmetics generally. Several of the ingredients here have genuinely good evidence for what was actually measured. It is a reason to read the promise and the evidence as two separate statements, because that is what they are.
Every grade in this analysis is published in machine-readable form, with the reasoning for each row, under CC BY 4.0.
The dataset also carries the wider register — grades by outcome, citation identifiers, and legal status by region — so the claim rows can be checked against the underlying evidence rather than taken on our word.
Suggested citation: Bilenko, Y. (2026). Ingredient Evidence Register (Version 1.2.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.21474487
Neutral evidence-based reference. Not medical advice, diagnosis, or treatment. Vallydia sells cosmetic serums; the commercial interest is disclosed on every page carrying a grade.
A credentialed reviewer (PharmD / PhD / MD) will be named before this entry is finalised. Until then, treat it as a working draft. Last updated 2026-07-21.
Related reading: Palmitoyl Pentapeptide-4 · how we grade.
A neutral reference and a lawful-lane shop. Registered in Spain. Information for those who seek it — never promotion.
This site provides neutral scientific reference and sells only products lawful in your region. Nothing here is medical advice, a recommendation, or an offer to supply unapproved medicines. No dosing or administration is published for research compounds. Cosmetic peptides per Regulation (EC) 1223/2009. Unapproved injectable peptides are neither sold nor advertised in the EU (Directive 2001/83/EC, Title VIII). © 2026 Vallydia SL — Registered in Spain.