GLP-1 & Incretin Science

Why Biopsy Endpoints Make Liver Trials Hard

JMWritten & reviewed by Jack Muncaster · Founder, UK PeptidesLast reviewed 2026-08-233 cited sources

MASH trials measure histological change on liver biopsy — resolution of steatohepatitis and improvement in fibrosis. These are scored by pathologists, carry documented inter-observer variability, and are surrogates for outcomes no trial can follow to completion.

Key facts

Endpoint source
Liver biopsy
Scored by
Pathologist judgement
Known issue
Inter-observer variability
Mitigation
Central reading, blinded
Endpoint status
Surrogate, not clinical outcome
Sampling issue
Biopsy samples a fraction of the organ

What is actually being measured

Two composite histological endpoints the field has converged on: resolution of steatohepatitis without worsening of fibrosis, and improvement in fibrosis without worsening of steatohepatitis. Both combine several scored features into a binary outcome, which is a considerable amount of judgement compressed into a yes or no.

The variability problem

Two competent pathologists reading the same slide do not always agree, and the same pathologist reading it twice does not always agree with themselves. That variability is well documented and it adds noise to the endpoint, which reduces a trial's power to detect a real treatment effect and can obscure modest benefits.

Research material referenced

Retatrutide 10mg — third-party HPLC tested

View — £59.99

Why sampling makes it worse

A needle biopsy retrieves a small fraction of a large organ, and MASH is not uniformly distributed through the liver. Two samples from the same patient can score differently because the disease differs between sites, not because anything changed. This is a limitation of the measurement rather than of the analysis.

How trials manage it

Central reading by a small panel, blinded to treatment assignment and often to sequence, with paired biopsies read together. That controls the variability rather than eliminating it, and it is why the reading methodology in a MASH trial deserves attention — a well-run central read is a substantive part of the evidence.

Why they are surrogates at all

What matters clinically is cirrhosis, liver failure, transplant and death, which take decades. No trial can wait. Histology is used because it lies on the path to those outcomes and can be measured in a few years — a reasonable compromise that remains a compromise.

The consequence for approvals

Regulators grant accelerated approvals on surrogate endpoints with confirmatory outcome evidence required afterwards. So a MASH approval is a conditional statement about a probable benefit, not a settled one — and that distinction almost never survives into coverage of these decisions.

Extended research context

The GLP-1 & Incretin Science deep dive

Deep dive: the two routes to a bigger effect

Every compound trying to beat GLP-1 alone has taken one of two routes. The first adds more receptors from the same hormone family — GIP in tirzepatide, GIP and glucagon in retatrutide. The second adds a non-incretin satiety hormone, which in practice means amylin: CagriSema combines cagrilintide with semaglutide, and amycretin engages both receptors from one molecule. Both routes work, because they recruit signalling pathways that do not fully overlap. Neither has escaped the constraint that binds all of them, which is that gastrointestinal tolerability worsens as effect size grows.

Deep dive: why a percentage is not a result

The most-quoted numbers in this field are the least comparable. REDEFINE 1 reported 22.7% and 20.4% for the same compound in the same trial — the first among participants who adhered to treatment, the second across everyone randomised. TRIUMPH-1 reported 28.3% in an uncomplicated obesity population while TRIUMPH-3 reported up to 22.6% in adults with established cardiovascular disease, using the same compound. Before any two figures can be compared they have to match on estimand, population, duration, comparator and whether the number is placebo-adjusted. Most published comparisons match on none of them.

Deep dive: what happens after the trial stops

Every headline figure describes weight while treatment continues. The STEP-1 extension found that a year after semaglutide was stopped, participants had given back roughly two-thirds of what they lost, moving from 17.3% mean reduction to a net 5.6% — though average weight remained below baseline and nearly half stayed at least 5% down. Meta-analysis puts regain at around 0.8 kg per month. This is why maintenance studies such as TRIUMPH-6 matter more to the field's future than another two points of peak reduction.

Research applications

  • Comparing incretin and amylin compounds on a like-for-like basis
  • Interpreting estimands, thresholds and placebo-adjusted figures in trial reports
  • Tracking the obesity pipeline across sponsors and jurisdictions
  • Understanding receptor pharmacology behind GLP-1, GIP, glucagon and amylin
  • Distinguishing licensed medicines from investigational compounds

Handling checklist

  • Identify which estimand a quoted percentage comes from before citing it
  • Check the trial population and baseline BMI against the comparison you are making
  • Confirm the duration and whether the reduction curve had plateaued
  • Read discontinuation rates alongside efficacy figures
  • Verify every NCT identifier against ClinicalTrials.gov rather than secondary reporting

Common research-handling mistakes

Learnt from thousands of researcher orders across our UK labs.

Comparing headline percentages across different trials

Fix: Population, duration, estimand and comparator all differ; the numbers are not interchangeable.

Quoting the larger of two figures from the same trial

Fix: Name the estimand. Efficacy and treatment-policy answer different questions.

Treating peak reduction as a durable outcome

Fix: Substantial regain follows cessation across the class; peak figures describe a maintained state.

Assuming an oral route means a weaker mechanism

Fix: Route and receptor count are independent. Orforglipron is weaker because it hits one receptor, not because it is a tablet.

Reading investigational compounds as available treatments

Fix: Most of this pipeline holds no authorisation anywhere; mazdutide is approved only in China.

Continue researching

Peer-reviewed guides, comparators and matched reference materials.

Related questions researchers ask

  • Which weight-loss compound produces the largest reduction?
  • What is the difference between CagriSema and amycretin?
  • What is an amylin receptor agonist?
  • How much weight is regained after stopping a GLP-1?
  • Why does CagriSema report two different percentages?
  • Why is orforglipron less effective than retatrutide?

Frequently asked questions

Why do MASH trials need biopsies?
Because the endpoints are histological — inflammation, ballooning and fibrosis scored on tissue. No blood test substitutes for them at present.
How reliable is histological scoring?
It carries documented inter-observer variability, and biopsy sampling captures only a fraction of a non-uniformly affected organ.
Why are these endpoints accepted?
Because the real outcomes — cirrhosis, liver failure, death — take decades, and histology lies on the path to them and can be measured in years.

Primary sources & clinical trials

Peer-reviewed research and registered trials from PubMed, ClinicalTrials.gov, PubChem, FDA and NIH. All links open in a new tab and point to the primary source, so every claim can be verified at origin.

JM

Written and reviewed by

Jack Muncaster · Founder, UK Peptides

Jack founded UK Peptides in Manchester after repeatedly receiving research compounds with missing or recycled paperwork. He is responsible for supplier selection, batch release decisions and the content published in this research library. Every article here is sourced to primary literature and every product page to a signed third-party certificate.

More GLP-1 & Incretin Science articles

Popular across the research hub

One flagship guide from every other research category — keep exploring.

Research use only. The information above is provided for scientific and educational reference. Compounds referenced are not approved for human use and are supplied for in vitro research or reference-material purposes only. No efficacy, safety, or therapeutic claims are made.