How to read a supplement study without a science degree
You do not need statistics to read a supplement study well. You need to ask five ordinary questions in a fixed order, and the first one — who was this actually done in — quietly disqualifies most of what you will ever see in a headline.
Start with the species, not the result
Three very different things get reported using the same word, study.
- In vitro. Cells in a dish, or an enzyme in a test tube. Useful for working out how something might behave. Says nothing about whether it survives your stomach, reaches your bloodstream, or does anything once it gets there.
- Animal. Usually rodents, often at doses that would be enormous scaled to a human body, often given by injection rather than by mouth.
- Human. People, taking the thing, in a form and an amount a person could plausibly take.
Press releases rarely lead with which one it was. A compound that does something dramatic to cells at a concentration you could never reach by swallowing it is a lead for researchers, not a reason to buy anything. When you find the study behind a claim, the answer to this question is usually in the first line of the abstract, and it is often the end of the conversation.
Evidence grade: well established that findings in cells and in rodents frequently fail to reproduce in humans.
How many people, and for how long
Small studies are not fraudulent. They are noisy. With twenty participants, ordinary chance can produce a large-looking effect in either direction, and the studies that happen to land on a large effect are precisely the ones that get written up, tweeted, and turned into product copy. That is why an n=20 result often travels further than a careful trial of eight hundred people that found very little.
Two habits help. Look for the number of participants before you look at the result. Then look at how long they were followed: four weeks tells you something different from four days, and almost nothing about a year.
Evidence grade: well established that small samples produce unstable, exaggerated effect estimates.
Randomised and blinded, or simply observed
In a randomised controlled trial, a coin flip decides who gets the real thing and who gets the placebo, and ideally neither the participants nor the researchers know which is which until the end. That structure exists for one reason: to stop expectation and selection from doing the work the ingredient was supposed to do.

Observational studies follow people who already made their own choices. They are valuable, sometimes the only ethical option, and they cannot separate the supplement from the person who takes it. People who take supplements also tend to sleep more, exercise more, and see a doctor more. That bundle is hard to untangle, and the tidiest observational finding still cannot tell you which strand mattered.
Evidence grade: well established that randomisation and blinding reduce bias relative to observational designs.

What did they measure, and would you have noticed it
This is the question that separates a real finding from a technically true one. Many trials measure a surrogate endpoint — a marker in the blood, a score on a questionnaire, a millisecond difference on a computerised attention task — because markers are cheap, fast, and move reliably.
A marker moving is not the same as your Tuesday afternoon going better. Sometimes the two are tightly linked. Often nobody has checked. When you read a claim, ask what was actually recorded, and whether a person living that life would have felt the difference described.
Who paid for it
Industry funding does not make a study wrong, and a great deal of good nutrition research is funded by the companies with a reason to run it. It does shift the odds. Across many fields, industry-sponsored trials report favourable results more often than independently funded ones, through mechanisms that are mostly mundane: which comparison was chosen, which endpoint was declared primary, whether a disappointing result was published at all.
Funding is disclosed at the bottom of almost every paper. Read it, note it, and weight the finding a little more cautiously — do not throw it away.
Evidence grade: some evidence for a consistent sponsorship effect on reported outcomes.
The studies you never saw
Publication bias is the most important idea on this page and the least visible. Journals have historically preferred striking results, and researchers know it. A trial that found nothing is more likely to sit in a drawer than to appear in print.
The consequence is structural: the published literature on any popular ingredient is skewed toward the results that worked. You are reading the winners of a competition you cannot see the entrants of. Trial registries were created partly to fix this, which is why a study that was registered before it started is worth more than one that was not.
Evidence grade: well established that null results are systematically under-published.
Meta-analysis: a real tool with a garbage-in problem
A meta-analysis pools many trials into one estimate. Done well it is the closest thing the field has to a verdict. Done badly it is an average of weak studies, presented with the authority of arithmetic.
Pooling twelve small, unblinded, industry-funded trials does not produce one large, rigorous trial. It produces a confident-looking number built from twelve shaky ones, and if publication bias filtered the inputs, the output inherits the skew. Good meta-analyses say so — they assess the quality of what they included and test how much the answer changes when the weakest studies are removed. If a meta-analysis does not discuss the quality of its inputs, treat the headline number as a starting point.

Significant is not the same as noticeable
Statistical significance answers one narrow question: how likely is a result this large if the thing did nothing at all. It says nothing about size. With enough participants, a genuinely tiny difference becomes statistically significant.
So look past the p-value to the effect size and the range around it. A sleep trial can report that participants fell asleep four minutes faster, significant across nine hundred people, and four minutes is not something you would ever detect in your own bed. Both statements are true at once. The honest version of the claim is the second one.
Five questions, applied to any headline
- In whom? Cells, animals, or people.
- How many, and for how long? A number and a duration, before the result.
- Randomised and blinded, or observed? And was it registered in advance.
- What was measured? A marker, or something you would have noticed.
- Who funded it, and what is missing? Including the trials that were never published.
Apply these to us. If we ever write something on this site that cannot survive all five, we have written it badly, and you should tell us. Our companion page on how we write about evidence explains the grading language we use, and how to read a Supplement Facts panel covers the other half of the job.

What this cannot do for you
Reading studies well makes you harder to sell to. It does not make you a scientist, and it will not tell you whether a given thing works in your particular body — trials report averages, and an average can hide people who improved a lot and people who got slightly worse.
It also cannot detect fabricated data, catch an ingredient that was mislabelled before it reached the bottle, or account for a medication you already take. If you are managing an ongoing health issue, are pregnant or nursing, or take prescription medication, that conversation belongs with a doctor, not with an abstract.
Where our guide fits
We wrote the CALMÉA Focus Kit because doing this properly for every ingredient takes hours nobody has. It is an 83-page illustrated PDF in three parts — Train, Rest, Fuel — across fifteen chapters, and every claim in it carries a grade for how strong the evidence behind it actually is. Where the honest answer is early evidence, it says early evidence. Where the honest answer is that the research has not settled, it says that too.
It is a reading guide, not a treatment plan, and it is deliberately duller than the headlines it exists to defuse.
FAQ
Does a study in mice mean anything for humans?
It means the question is worth asking in humans. Dose, route of administration and physiology all differ, and a large share of promising animal findings do not reproduce in human trials.
How many participants is enough?
There is no single number, because it depends on how large the effect is expected to be. As a rough habit: treat anything under about fifty participants as preliminary, and be sceptical of any dramatic claim built on a single small trial.
Is a placebo-controlled trial always better than an observational study?
For cause and effect, generally yes. Observational work is better for long-term patterns and rare outcomes that a trial could never ethically or practically capture. They answer different questions.
What does it mean when a study is preregistered?
The researchers published their plan — outcomes, sample size, analysis — before collecting data. It makes it much harder to quietly switch to whichever result turned out best, so preregistered trials deserve more weight.
Should I ignore any study funded by a supplement company?
No. Weight it more cautiously, check whether it was preregistered and independently analysed, and look for whether anyone unconnected to the company has replicated it.
Content is for general educational purposes only and is not medical advice.
These statements have not been evaluated by the Food and Drug Administration. This product is not intended to diagnose, treat, cure, or prevent any disease.