GLP-1 & Incretin Science

A p-Value of 0.0035 and an Effect of 0.16

JMWritten & reviewed by Jack Muncaster · Founder, UK PeptidesLast reviewed 2026-08-243 cited sources

Statistical significance measures how confidently a difference can be distinguished from chance. It says nothing at all about how large that difference is. In a large trial the two routinely come apart, and a robustly significant result can still be clinically marginal.

Key facts

What a p-value measures
Compatibility of the data with no difference
What it does not measure
The size or importance of a difference
What determines detectable size
Sample size and measurement precision
Worked example
REIMAGINE 2, HbA1c, 603 vs 605 participants
Result
-1.91 vs -1.75 percentage points
Difference
-0.16 points (95% CI -0.27 to -0.05)
p-value
0.0035
What to read instead
The effect size with its confidence interval

The confusion this article is about

A statistically significant result is routinely reported as though significance were a measure of magnitude - as though a smaller p-value meant a bigger effect. It does not. A p-value answers one question: how compatible is the observed data with there being no difference at all. A very small p-value says a difference is very probably real. It is silent on whether that difference is large enough for anyone to care about, and the two questions can point in opposite directions in the same result.

Why large trials make this worse

The size of difference a trial can detect shrinks as the trial grows. Double the participants and the detectable difference falls; multiply them tenfold and it falls further. A trial with six hundred participants per arm measuring a precise laboratory endpoint can distinguish differences from zero that are far below anything that would change a decision. This is not a flaw. It is the intended behaviour of a well-powered trial, and it means that in a large study, statistical significance becomes a weak filter and the effect size does all the work.

The worked example

REIMAGINE 2 compared a fixed combination against one of its own components at matching dose, with 603 and 605 participants in the two arms, on the change in glycated haemoglobin over 68 weeks. HbA1c fell 1.91 percentage points on the combination and 1.75 on the component. The difference was 0.16 percentage points, the 95% confidence interval ran from 0.27 to 0.05, and the p-value was 0.0035. Every one of those numbers is correctly reported and the conclusion of superiority is correct. A reader who saw only 'superior, p=0.0035' would form an impression of a substantially better drug, and the effect size does not support that impression.

Read the interval, not the point estimate alone

The interval here is the most informative single object in the result. It runs from 0.27 to 0.05 percentage points, so the data are compatible with a difference as small as 0.05. Because the entire interval sits above zero, the difference is established. Because the entire interval sits well below the thresholds usually treated as clinically important for this measure, the magnitude is established too - as modest. An interval that excludes zero and also excludes any large effect is telling you two things at once, and both are worth reporting.

The mirror image, which this library has also covered

The opposite failure is a point estimate with a very wide interval, where a striking-looking number is compatible with almost anything - a published meta-analysis on this class reported a risk ratio of 0.568 with an interval from 0.077 to 4.205. Between them these two cases define the habit worth forming: a point estimate on its own is never enough. A narrow interval near zero says a small effect is real. A wide interval says nothing was measured. Only the interval distinguishes them, and only the interval tells you which situation you are in.

How to state a result like this honestly

By reporting both halves. The combination was superior to the component on this endpoint, and the difference was 0.16 percentage points, with a confidence interval from 0.27 to 0.05. That sentence is longer than the headline version and it is the only version that leaves a reader with the impression the data actually support. Nothing in this article concerns any product supplied on this site, and the medicines referred to are licensed prescription products that research material is not an alternative to.

Extended research context

The GLP-1 & Incretin Science deep dive

Deep dive: the two routes to a bigger effect

Every compound trying to beat GLP-1 alone has taken one of two routes. The first adds more receptors from the same hormone family — GIP in tirzepatide, GIP and glucagon in retatrutide. The second adds a non-incretin satiety hormone, which in practice means amylin: CagriSema combines cagrilintide with semaglutide, and amycretin engages both receptors from one molecule. Both routes work, because they recruit signalling pathways that do not fully overlap. Neither has escaped the constraint that binds all of them, which is that gastrointestinal tolerability worsens as effect size grows.

Deep dive: why a percentage is not a result

The most-quoted numbers in this field are the least comparable. REDEFINE 1 reported 22.7% and 20.4% for the same compound in the same trial — the first among participants who adhered to treatment, the second across everyone randomised. TRIUMPH-1 reported 28.3% in an uncomplicated obesity population while TRIUMPH-3 reported up to 22.6% in adults with established cardiovascular disease, using the same compound. Before any two figures can be compared they have to match on estimand, population, duration, comparator and whether the number is placebo-adjusted. Most published comparisons match on none of them.

Deep dive: what happens after the trial stops

Every headline figure describes weight while treatment continues. The STEP-1 extension found that a year after semaglutide was stopped, participants had given back roughly two-thirds of what they lost, moving from 17.3% mean reduction to a net 5.6% — though average weight remained below baseline and nearly half stayed at least 5% down. Meta-analysis puts regain at around 0.8 kg per month. This is why maintenance studies such as TRIUMPH-6 matter more to the field's future than another two points of peak reduction.

Research applications

  • Comparing incretin and amylin compounds on a like-for-like basis
  • Interpreting estimands, thresholds and placebo-adjusted figures in trial reports
  • Tracking the obesity pipeline across sponsors and jurisdictions
  • Understanding receptor pharmacology behind GLP-1, GIP, glucagon and amylin
  • Distinguishing licensed medicines from investigational compounds

Handling checklist

  • Identify which estimand a quoted percentage comes from before citing it
  • Check the trial population and baseline BMI against the comparison you are making
  • Confirm the duration and whether the reduction curve had plateaued
  • Read discontinuation rates alongside efficacy figures
  • Verify every NCT identifier against ClinicalTrials.gov rather than secondary reporting

Common research-handling mistakes

Learnt from thousands of researcher orders across our UK labs.

Comparing headline percentages across different trials

Fix: Population, duration, estimand and comparator all differ; the numbers are not interchangeable.

Quoting the larger of two figures from the same trial

Fix: Name the estimand. Efficacy and treatment-policy answer different questions.

Treating peak reduction as a durable outcome

Fix: Substantial regain follows cessation across the class; peak figures describe a maintained state.

Assuming an oral route means a weaker mechanism

Fix: Route and receptor count are independent. Orforglipron is weaker because it hits one receptor, not because it is a tablet.

Reading investigational compounds as available treatments

Fix: Most of this pipeline holds no authorisation anywhere; mazdutide is approved only in China.

Continue researching

Peer-reviewed guides, comparators and matched reference materials.

Related questions researchers ask

  • Which weight-loss compound produces the largest reduction?
  • What is the difference between CagriSema and amycretin?
  • What is an amylin receptor agonist?
  • How much weight is regained after stopping a GLP-1?
  • Why does CagriSema report two different percentages?
  • Why is orforglipron less effective than retatrutide?

Frequently asked questions

Is a small effect worthless?
Not necessarily. A small average effect across a very large population can matter at population level, and a small mean can hide a subgroup with a large one. What is not defensible is presenting a small effect as though it were large because the p-value is small.
What counts as a clinically important difference in HbA1c?
There is no single agreed number and it depends on baseline and context, but thresholds discussed in the literature are generally several times larger than 0.16 percentage points. The safest approach is to compare the observed effect against what established agents achieve rather than against a fixed rule.
Should trials stop reporting p-values?
They serve a purpose - distinguishing a real difference from noise is a genuine question. The argument is about prominence: reporting the p-value first and the effect size second inverts their importance for a reader trying to decide whether a result matters.
How do I spot this in coverage?
Look for the word 'significant' unaccompanied by a number. If a summary states superiority without giving the size of the difference and its interval, it has reported the less informative half of the result.

Primary sources & clinical trials

Peer-reviewed research and registered trials from PubMed, ClinicalTrials.gov, PubChem, FDA and NIH. All links open in a new tab and point to the primary source, so every claim can be verified at origin.

JM

Written and reviewed by

Jack Muncaster · Founder, UK Peptides

Jack founded UK Peptides in Manchester after repeatedly receiving research compounds with missing or recycled paperwork. He is responsible for supplier selection, batch release decisions and the content published in this research library. Every article here is sourced to primary literature and every product page to a signed third-party certificate.

More GLP-1 & Incretin Science articles

Popular across the research hub

One flagship guide from every other research category — keep exploring.

Research use only. The information above is provided for scientific and educational reference. Compounds referenced are not approved for human use and are supplied for in vitro research or reference-material purposes only. No efficacy, safety, or therapeutic claims are made.