Skip to content

Claim Forensics

What "Up To" Actually Means on a Supplement Label

The average is what most people got. The maximum is what one person got. Guess which one ends up on the label.

A magnifying glass with a brass rim resting on a page of tiny grey fine print beside a plain white supplement bottle on dark slate, lit from one side.

Almost every weight-loss supplement page has a number on it with two small words in front. Up to 12 pounds. Up to 4 inches. Up to 30% more fat burned.

Those two words do an enormous amount of work. They take the single best result anybody in the study got and put it where your eye lands first.

This isn't a takedown of one brand. It's a walkthrough of a technique that runs through the whole category, and it takes about 90 seconds to defuse once you know which pieces to pull apart.

A ceiling is not a prediction

An average tells you what a typical participant experienced. A maximum tells you what happened to one person, once, in one group. Both are true numbers. Only one of them is a reasonable thing for you to plan around.

"Up to" quietly swaps the second for the first and lets you fill in the rest.

Here is a live example we verified on the brand's own product page on 28 August 2026. Evolv GLP-1 leads with "12+ lbs lost, up to, in 8 weeks" and "4+ inches off the waist, up to, in 8 weeks." The dagger footnote on the same page reads, in full caps: "BASED ON 8-WEEK PER-PROTOCOL RESULTS FROM THE EXPERIMENTAL GROUP OF A RANDOMIZED, CONTROLLED CLINICAL STUDY. AT WEEK 8, UP TO 12+ LBS OF WEIGHT LOSS AND UP TO 4+ INCHES OF WAIST REDUCTION WERE OBSERVED."

Read that footnote slowly and it tells you three separate things: the figure is a maximum, it comes only from the group that got the product, and it comes only from the people who completed the protocol. The group averages the company published in its own interim write-up are 2.87% body weight and 4.02% waist versus control at week 8. On a 200-pound person, 2.87% is under 6 pounds. Our full teardown of that claim set goes further into the arithmetic.

Per-protocol is the friendlier half of the room

"Per-protocol" means the analysis counted the people who finished the study the way they were supposed to. Anyone who quit, skipped doses, or dropped out gets left out of the math.

That sounds reasonable and it is not neutral. People who quit a supplement trial often quit because it wasn't working for them, or because it made them feel bad. Removing them removes the worst results.

How much does it matter? A meta-epidemiological study in The BMJ looked at 167 randomized trials across 14 meta-analyses of osteoarthritis treatments. Only 39 of them (23%) analyzed every patient they randomized. In the other 128, effect sizes leaned more favorable, and the gap widened in exactly the settings you'd worry about: trials with large claimed benefits, and trials of complementary medicine. When the authors restricted the analyses to trials that excluded nobody, the estimated benefits shrank and the p-values got bigger (Nüesch et al., BMJ 2009).

Their conclusion is worth holding onto: exclusions bias the result, and the direction and size of that bias are unpredictable. You can't correct for it in your head. You can only notice it's there.

The counterpart is intention-to-treat, which counts everyone who was randomized, quitters included. It's a worse number for marketing and a better number for you, because you're about to be a person who might quit.

The number that was never measured

Some headline figures aren't observations at all. They're arithmetic performed on another headline figure.

Evolv's third banner claim is "~750 fewer calories consumed per day without restrictive dieting." Nobody weighed anybody's food. The footnote says so: "A 12-LB REDUCTION OVER 8 WEEKS REFLECTS AN ESTIMATED ENERGY DEFICIT EQUIVALENT TO ~750 CALORIES PER DAY."

That's the old 3,500-calories-per-pound rule, run backwards, starting from the maximum rather than the average. 12 times 3,500, divided by 56 days, gives 750.

The rule itself has been out of date for 15 years. Kevin Hall's modelling work in The Lancet showed that bodyweight responds to a change in energy intake slowly, with a half-time of roughly a year, because expenditure adapts downward as weight falls. Static arithmetic overshoots (Hall et al., 2011). So the calorie claim is a stale formula applied to a best case, presented as a measured behavior change.

The primary outcome that moved

Trials declare, in advance, the one thing they're testing. That's the primary outcome. It exists so nobody can measure 30 things, find the 2 that worked, and write the press release about those.

People do it anyway. A cohort study in JAMA compared the protocols of 102 randomized trials against what actually got published. Half of the efficacy outcomes and 65% of the harm outcomes per trial were reported incompletely. Statistically significant results were about 2.4 times more likely to be fully reported than non-significant ones for efficacy, and 4.7 times more likely for harms. And 62% of the trials had at least one primary outcome that was changed, introduced, or quietly dropped between protocol and publication (Chan et al., JAMA 2004).

The best detail in that paper: 86% of the authors who responded to a survey denied there were any unreported outcomes, while the protocols sat there saying otherwise.

And when the main result comes back empty

Researchers call the salvage operation "spin." It has recognizable moves.

An analysis of 110 randomized trials with clearly non-significant primary outcomes, published across 10 high-impact surgical journals, found spin in 40% of abstracts and 35% of main texts. 25 of those papers (23%) recommended the intervention anyway (Arunachalam et al., Ann Surg 2017). A separate systematic review of 92 plastic surgery trials with non-significant primary outcomes found spin strategies in 85% of them, the commonest being to reframe a failed difference as evidence of equivalence (Yuan et al., Plast Reconstr Surg 2023).

Those are peer-reviewed medical journals with editors and reviewers. A supplement product page has neither.

"Statistically significant" is not the same as "worth your money"

Even a clean, honest, fully reported average can be too small to feel.

A systematic review in the International Journal of Obesity pooled 67 randomized placebo-controlled trials of dietary supplements containing isolated organic compounds. Three ingredients beat placebo to statistical significance: chitosan at 1.84 kg, glucomannan at 1.27 kg, and conjugated linoleic acid at 1.08 kg. The authors had set 2.5 kg as their threshold for a clinically meaningful difference. None of the three cleared it. Their verdict was that there's currently not enough evidence to recommend any of these supplements for weight loss (Bessell et al., Int J Obes 2021).

Glucomannan is the fiber in Lipozene, and we've written up the wider evidence base in our glucomannan deep dive. A 1.27 kg average over the course of a trial is a real finding. It is also roughly the amount your weight swings across a normal week.

What the FTC actually asks for

None of this is a legal grey zone, and the agency does act on it: it once banned a manufacturer outright over the claims behind its product, which we cover in the evidence on appetite control patches. The Federal Trade Commission published updated Health Products Compliance Guidance in December 2022, its first full rewrite of the advice in nearly 25 years, and it addresses this pattern head on.

On dramatic results, the guidance says testimonials reporting outcomes more dramatic than users can generally expect are likely to be deceptive, and that trying to disclaim them with lines like "Results not typical" doesn't cure the problem. Its worked example is a weight-loss ad quoting a woman who lost 16 pounds in 8 weeks, with fine print reading "These results are not typical." Her experience is real. The trial behind the product showed an average of 4 pounds over placebo in the same 8 weeks. The FTC's position is that the disclaimer fails, and that the fix is stating the averages for both groups, next to the quote, in prominent type.

The guidance is equally blunt about after-the-fact analysis. It warns that a post hoc analysis departing from the original protocol can indicate researchers are data mining or p-hacking to find something positive in a study that otherwise failed, and gives an example in which selectively promoting a favorable subgroup result, when the trial overall found nothing, is itself deceptive.

There's also a quieter one worth knowing. If the trial participants were all following a reduced-calorie diet and exercising, the ad has to say so, because otherwise the pill is taking credit for the diet.

Five questions that take about 90 seconds

  • Is that number an average or a maximum? If the words "up to" appear, assume maximum until the page tells you otherwise.
  • Compared to what? A weight change in the treatment group alone is not a result. The placebo group also loses weight. The only number that means anything is the difference between them.
  • Who got counted? Look for per-protocol versus intention-to-treat, and for how many people started versus finished.
  • Was it measured or calculated? Calorie and percentage figures are often derived from a weight figure, not observed.
  • Where is it published? A registered trial with a peer-reviewed write-up is a different object from an interim summary on the company's own blog.

Read the footnote first. In this category the footnote is usually accurate, and it's usually the only part of the page that is.

How we score and compare products across metabolism and weight is set out on our methodology page, and the head-to-head that put this post in motion is here.

References

  • Nüesch E, Trelle S, Reichenbach S, et al. The effects of excluding patients from the analysis in randomised controlled trials: meta-epidemiological study. BMJ. 2009;339:b3244. PMID 19736281
  • Chan AW, Hróbjartsson A, Haahr MT, Gøtzsche PC, Altman DG. Empirical evidence for selective reporting of outcomes in randomized trials: comparison of protocols to published articles. JAMA. 2004;291(20):2457-2465. PMID 15161896
  • Arunachalam L, Hunter IA, Killeen S. Reporting of randomized controlled trials with statistically nonsignificant primary outcomes published in high-impact surgical journals. Ann Surg. 2017;265(6):1141-1145. PMID 27257737
  • Yuan M, Wu J, Li A, et al. "Spin" in plastic surgery randomized controlled trials with statistically nonsignificant primary outcomes: a systematic review. Plast Reconstr Surg. 2023;151(3):506e-519e. PMID 36442055
  • Bessell E, Maunder A, Lauche R, Adams J, Sainsbury A, Fuller NR. Efficacy of dietary supplements containing isolated organic compounds for weight loss: a systematic review and meta-analysis of randomised placebo-controlled trials. Int J Obes (Lond). 2021;45(8):1631-1643. PMID 33976376
  • Hall KD, Sacks G, Chandramohan D, et al. Quantification of the effect of energy imbalance on bodyweight. Lancet. 2011;378(9793):826-837. PMID 21872751
  • Federal Trade Commission. Health Products Compliance Guidance. December 2022.

Research retrieved via PubMed. Brand claims and footnote text were read from the company's own product page on 28 August 2026. Nothing here is medical advice.

FAQ

Does "up to" mean the claim is false?

No, and that's what makes it effective. The number is usually real. It's the maximum result observed in the treatment group, which is a true measurement of one person's outcome. It just isn't a forecast of yours. The figure you want is the average difference between the treatment group and the placebo group, and that is almost never the number on the banner.

What is the difference between per-protocol and intention-to-treat?

Per-protocol counts only the participants who completed the study as instructed. Intention-to-treat counts everyone who was randomized, including people who dropped out or stopped taking the product. Per-protocol usually produces a bigger effect, because the people who quit are often the ones it wasn't working for. A BMJ meta-epidemiological study of 167 trials found that trials excluding patients from the analysis reported more favorable effects, and that restricting to trials with no exclusions shrank the estimated benefits (Nüesch et al., BMJ 2009).

Why is a calorie figure on a supplement page a warning sign?

Because intake is hard to measure and easy to calculate. A claim like "750 fewer calories a day" is often derived from a weight-loss figure using the old 3,500-calories-per-pound rule rather than observed from food records. Kevin Hall's modelling in The Lancet showed that rule overestimates, because bodyweight responds slowly to a change in intake and energy expenditure adapts as weight falls (Hall et al., 2011). Check the footnote for the word "estimated" or "equivalent to."

Is a statistically significant weight-loss result always meaningful?

Not necessarily. A meta-analysis of 67 placebo-controlled supplement trials found statistically significant differences for chitosan (1.84 kg), glucomannan (1.27 kg) and conjugated linoleic acid (1.08 kg), but none of them reached the authors' 2.5 kg threshold for clinical significance, and the review concluded there was insufficient evidence to recommend any of them for weight loss (Bessell et al., Int J Obes 2021). Statistical significance says an effect probably isn't chance. It says nothing about whether the effect is large enough to notice.

Products mentioned

#1

Calocurb GLP-1 Activator

Calocurb · capsule
16/25
Limited value

The best-evidenced single-mechanism appetite trigger — clinically interesting, but premium-priced and a pre-meal, short-window tool rather than all-day support.

Strength

Amarasate is one of the most directly studied appetite ingredients — multiple human RCTs, dosed on-label at 250 mg.

Watch-out

Acute, short-window effect: a pre-meal tool (up to 4 capsules/day), not all-day craving support.

Full breakdown →
#2

PhenQ

Wolfson Brands · capsule
Best for Daytime Energy
13/25
Fair value

Best for Daytime Energy: a caffeinated daytime option for some — but useless for after-dinner cravings, and it hides its doses.

Strength

Caffeine provides a genuine daytime energy and appetite-blunting kick.

Watch-out

Key actives are hidden in proprietary blends (a-Lacys Reset, Capsimax).

Full breakdown →
#3

Evolv GLP-1

Evolv · tablet
12/25
Limited value

A genuinely novel idea wrapped in the category’s worst claim hygiene: "up to" headlines several times the actual averages, an unpublished and unpowered interim readout, the highest price in our set, and no refunds.

Strength

Simple, caffeine-free two-tablet daily routine with no timing restriction — usable in the evening craving window.

Watch-out

The most expensive product in our set at about $4.93 a serving, and the only one with no money-back guarantee of any kind.

Full breakdown →
#4

Lipozene

Obesity Research Institute · capsule
12/25
Limited value

Cheap single-ingredient glucomannan — fine as a fiber, but it does only one thing.

Strength

Single, disclosed ingredient.

Watch-out

One pathway only (fiber satiety).

Full breakdown →