GLP-1 & Incretin Science
Failing to Show Non-Inferiority Is Not Showing Inferiority
A non-inferiority trial asks whether a new treatment is not worse than a comparator by more than a pre-specified margin. Failing to demonstrate that is a statement about where a confidence interval fell — it is not the same as demonstrating the treatment is inferior.
Key facts
- Question asked
- Is it not worse by more than margin X?
- Set in advance
- The margin, before any data
- Decided by
- Where the confidence interval falls
- Failing means
- Non-inferiority was not demonstrated
- Failing does not mean
- Inferiority was demonstrated
- Live example
- REDEFINE 4 (NCT06131437)
Why the design exists at all
Superiority trials ask whether a treatment beats a comparator. Sometimes that is the wrong question — a new option may be worth having because it is easier to take, better tolerated or cheaper, provided it is not meaningfully less effective. Non-inferiority formalises not meaningfully less effective into a number agreed before the trial runs.
The margin is the whole design
A non-inferiority margin is the largest difference that would still be considered acceptable. It has to be justified and fixed in advance, because a margin chosen after seeing the data could be set wherever it produces the desired conclusion. Everything about how such a trial reads depends on that number, and it is the first thing to look for.
Research material referenced
Retatrutide 10mg — third-party HPLC tested
What the result actually is
A confidence interval around the difference between treatments. If the whole interval sits within the margin, non-inferiority is demonstrated. If part of it falls outside, it is not. That is a statement about precision and position, not a verdict on the compound — and a trial can miss the threshold while still showing the treatment worked well.
The asymmetry most coverage misses
Three outcomes are possible, not two. Non-inferiority demonstrated. Non-inferiority not demonstrated. Or inferiority actually demonstrated, which requires the entire interval to sit on the wrong side. Reporting that lands on failed and therefore worse collapses the middle case into the third, and the middle case is where most failed non-inferiority trials land.
REDEFINE 4 as the live illustration
The ClinicalTrials.gov record for NCT06131437 states the primary outcome as confirming non-inferiority of CagriSema versus tirzepatide on relative change in body weight. That objective was not met. The reported figures were 23.0% against 25.5% on the efficacy estimand — a real difference, and also a result in which the smaller number is a large weight reduction.
How to read one properly
Find the pre-specified margin. Find the confidence interval. Check which of the three outcomes actually occurred. Then read the absolute numbers separately, because a compound can miss a comparative threshold and still be effective — those are two different questions and a headline usually answers neither.
Quick reference
| Where the confidence interval sits | Conclusion |
|---|---|
| Entirely within the margin | Non-inferiority demonstrated |
| Crosses the margin | Not demonstrated — verdict open |
| Entirely beyond the margin | Inferiority demonstrated |
Extended research context
The GLP-1 & Incretin Science deep dive
Deep dive: the two routes to a bigger effect
Every compound trying to beat GLP-1 alone has taken one of two routes. The first adds more receptors from the same hormone family — GIP in tirzepatide, GIP and glucagon in retatrutide. The second adds a non-incretin satiety hormone, which in practice means amylin: CagriSema combines cagrilintide with semaglutide, and amycretin engages both receptors from one molecule. Both routes work, because they recruit signalling pathways that do not fully overlap. Neither has escaped the constraint that binds all of them, which is that gastrointestinal tolerability worsens as effect size grows.
Deep dive: why a percentage is not a result
The most-quoted numbers in this field are the least comparable. REDEFINE 1 reported 22.7% and 20.4% for the same compound in the same trial — the first among participants who adhered to treatment, the second across everyone randomised. TRIUMPH-1 reported 28.3% in an uncomplicated obesity population while TRIUMPH-3 reported up to 22.6% in adults with established cardiovascular disease, using the same compound. Before any two figures can be compared they have to match on estimand, population, duration, comparator and whether the number is placebo-adjusted. Most published comparisons match on none of them.
Deep dive: what happens after the trial stops
Every headline figure describes weight while treatment continues. The STEP-1 extension found that a year after semaglutide was stopped, participants had given back roughly two-thirds of what they lost, moving from 17.3% mean reduction to a net 5.6% — though average weight remained below baseline and nearly half stayed at least 5% down. Meta-analysis puts regain at around 0.8 kg per month. This is why maintenance studies such as TRIUMPH-6 matter more to the field's future than another two points of peak reduction.
Research applications
- ▸Comparing incretin and amylin compounds on a like-for-like basis
- ▸Interpreting estimands, thresholds and placebo-adjusted figures in trial reports
- ▸Tracking the obesity pipeline across sponsors and jurisdictions
- ▸Understanding receptor pharmacology behind GLP-1, GIP, glucagon and amylin
- ▸Distinguishing licensed medicines from investigational compounds
Handling checklist
- ✓Identify which estimand a quoted percentage comes from before citing it
- ✓Check the trial population and baseline BMI against the comparison you are making
- ✓Confirm the duration and whether the reduction curve had plateaued
- ✓Read discontinuation rates alongside efficacy figures
- ✓Verify every NCT identifier against ClinicalTrials.gov rather than secondary reporting
Common research-handling mistakes
Learnt from thousands of researcher orders across our UK labs.
✗ Comparing headline percentages across different trials
Fix: Population, duration, estimand and comparator all differ; the numbers are not interchangeable.
✗ Quoting the larger of two figures from the same trial
Fix: Name the estimand. Efficacy and treatment-policy answer different questions.
✗ Treating peak reduction as a durable outcome
Fix: Substantial regain follows cessation across the class; peak figures describe a maintained state.
✗ Assuming an oral route means a weaker mechanism
Fix: Route and receptor count are independent. Orforglipron is weaker because it hits one receptor, not because it is a tablet.
✗ Reading investigational compounds as available treatments
Fix: Most of this pipeline holds no authorisation anywhere; mazdutide is approved only in China.
Continue researching
Peer-reviewed guides, comparators and matched reference materials.
Related questions researchers ask
- Which weight-loss compound produces the largest reduction?
- What is the difference between CagriSema and amycretin?
- What is an amylin receptor agonist?
- How much weight is regained after stopping a GLP-1?
- Why does CagriSema report two different percentages?
- Why is orforglipron less effective than retatrutide?
Frequently asked questions
- Does failing non-inferiority mean the drug is worse?
- No. It means non-inferiority was not demonstrated. Showing inferiority requires the entire confidence interval to fall beyond the margin.
- Why not just run a superiority trial?
- Because a new option can be worth having for tolerability, convenience or cost if it is not meaningfully less effective — which is what non-inferiority tests.
- What should I look for first?
- The pre-specified margin. It must be fixed before the data, and everything about how the result reads depends on it.
Primary sources & clinical trials
Peer-reviewed research and registered trials from PubMed, ClinicalTrials.gov, PubChem, FDA and NIH. All links open in a new tab and point to the primary source, so every claim can be verified at origin.
- TrialClinicalTrials.gov · CagriSema compared to tirzepatide, Phase 3 (NCT06131437)clinicaltrials.gov
- PubMedBoisseau W et al., Understanding non-inferiority trials — Neurochirurgie 2026 (PMID 41353860)pubmed.ncbi.nlm.nih.gov
- TrialClinicalTrials.gov · TRIUMPH-1 (NCT05929066) — Retatrutide pivotal obesity trialclinicaltrials.gov
- TrialClinicalTrials.gov · TRIUMPH-6 (NCT06859268) — Maintenance of weight reductionclinicaltrials.gov
- RefNovo Nordisk · CagriSema REDEFINE 1, published in NEJMprnewswire.com
- PubMedAmycretin phase 1b/2a subcutaneous study — PubMed (PMID 40550231)pubmed.ncbi.nlm.nih.gov
- RefTrajectory of weight regain after GLP-1 cessation — eClinicalMedicinethelancet.com
- PubMedOrforglipron: A Comprehensive Review — Int J Mol Sci 2026 (PMID 41683830)pubmed.ncbi.nlm.nih.gov
- EMAICH E9(R1) — estimands in clinical trials (EMA)ema.europa.eu
- GuidelineGoogle — Creating helpful, reliable, people-first contentdevelopers.google.com
Written and reviewed by
Jack Muncaster · Founder, UK Peptides
Jack founded UK Peptides in Manchester after repeatedly receiving research compounds with missing or recycled paperwork. He is responsible for supplier selection, batch release decisions and the content published in this research library. Every article here is sourced to primary literature and every product page to a signed third-party certificate.
More GLP-1 & Incretin Science articles
- Satiety and Nausea Run Through Different CircuitsA 2024 Nature paper showed hindbrain GLP-1 receptor circuits for satiety and aversion are dissociable. Nausea may not be the price of appetite suppression.
- Three Head-to-Heads, One Common ComparatorSURMOUNT-5, REDEFINE 4 and TRIUMPH-5 all use tirzepatide. What it means when one compound becomes the field's reference point.
- What Open-Label Does to a Weight-Loss TrialREDEFINE 4's registry record lists masking as none. Why an unblinded weight trial is weaker evidence, and where the effect creeps in.
- Two Answers to the Same QuestionCagriSema pairs GLP-1 with amylin; tirzepatide pairs it with GIP. REDEFINE 4 was effectively the first direct test between those strategies.
- Semaglutide and Liver Disease: What ESSENCE TestedA 1,205-participant Phase 3 in metabolic dysfunction-associated steatohepatitis, published in NEJM in June 2025. What it measured and why liver trials are hard.
Popular across the research hub
One flagship guide from every other research category — keep exploring.
- Retatrutide ResearchRetatrutide Molecular Structure and Properties
- GHK-Cu (Copper Peptide)GHK-Cu Mechanism: What Is Actually Proposed
- TB-500 (Thymosin β4 fragment)How Actin Sequestration Works
- BPC-157 (Pentadecapeptide)BPC-157 Structure and Sequence
- CJC-1295 & IpamorelinIpamorelin Sequence and Chemical Identity
- Peptide ReferenceLyophilisation: Why Peptides Arrive as Powder
- Bacteriostatic WaterBenzyl Alcohol Is Not an Inert Excipient
- Research & Regulatory NewsTRIUMPH-4: The First Retatrutide Phase 3 to Read Out
- MOTS-c (Mitochondrial Peptide)MOTS-c Storage, Stability and Reconstitution
- Semax (ACTH Fragment Peptide)Semax CAS Number and Chemical Identity
- Selank (Tuftsin Analogue)Selank Storage, Stability and Reconstitution
- DSIP (Delta Sleep-Inducing Peptide)An Antibody Reports What It Binds, Not What Is There
- KLOW (Blend)Four Identities, Four Masses, One Vial
- GLOW (Blend)One Volume, Three Peptides
- MT-2 (Melanotan II)A 1995 Paper, and a Structure That Checks Out Exactly
- IGF-1 LR3A Natural Experiment in Absent IGF-1
- GlutathioneGlutathione Proposed as a Carrier, Not Just a Buffer
- NAD+Eighty Daltons Apart, and Metabolically Opposed
- KPVFragment Logic: KPV Against Its Parent