This website uses cookies

Read our Privacy policy and Terms of use for more information.

Clinical Takeaway
Domain Key Finding Evidence Level
The mechanism Rodent C-cells are dense with GLP-1 receptors and proliferate on exposure. Primate and human C-cells do not respond: no adenylate cyclase activation, no calcitonin rise at 60× exposure for 20 months. Preclinical, species-divergent
The headline number No detectable excess, on estimates that are not precise. Pooled OR 1.37 (0.82–2.31) across 48 randomized trials; six-country cohort HR 0.81 (0.59–1.12); four-comparator claims analysis 0.78–1.03. Meta-analysis, n=94,245
The reversal The 2023 French signal tracks screening, not biology. Risk was elevated only in year one (HR 1.85), then vanished, while thyroid ultrasound ran 2.1% vs 1.5% in GLP-1 starters. Target trial emulation, n=351,000
The caveat Medullary thyroid cancer has never been testable: too few events in any cohort to analyze. Follow-up tops out near three years, and nearly all data come from type 2 diabetes, not obesity-only users. Untested, not negative
Where it fits Ask about personal and family MTC and MEN2 history at every consult. Personal MTC or MEN2 ends the discussion; a family history of apparently sporadic MTC opens shared decision making. Do not order calcitonin or screening ultrasound: ordering them is what generates the association. Practice unchanged

Orforglipron, the first oral non-peptide GLP-1 receptor agonist, does not activate the rat GLP-1 receptor. Investigators fed it to rats for two years at exposures up to 26 times the maximum recommended human dose and got no tumors. Its label, revised this July, opens with a boxed warning about rodent thyroid C-cell tumors anyway.

That is the state of the question in 2026. The warning has become a property of the receptor class rather than a finding about any particular drug, inherited by molecules that cannot even produce the effect it describes. Meanwhile the human evidence has been accumulating steadily, across six national health systems and 93 clinical trials, and the myth worth killing is not that the drugs are dangerous. It is the quieter assumption on the other side, that because the studies keep coming back null, the question was never serious. It was serious. The rodent data were real. What changed is that we finally have enough human data to say what they meant.

Where the warning came from

In 2010, Bjerre Knudsen and colleagues published the experiments that put the black box on the label. They gave liraglutide to rats and mice, and the animals' thyroid C-cells did exactly what a receptor-mediated effect should do: calcitonin release, upregulated calcitonin gene expression, C-cell hyperplasia, and eventually adenomas and carcinomas at clinically relevant exposures. The GLP-1 receptor was localized to the C-cells themselves. This was not a fluke of one compound or an artifact of absurd dosing.

The same paper also contained the argument against extrapolating it. In cynomolgus monkeys, GLP-1 receptor agonists did not activate adenylate cyclase and did not generate calcitonin release. Monkeys given liraglutide at roughly 60 times human exposure for 20 months showed no C-cell hyperplasia. And in human participants, mean calcitonin after two years of liraglutide sat at the low end of the normal range and stayed there.

So the boxed warning was, from the first day, a warning about a species. Rodents carry abundant C-cells dense with GLP-1 receptors. Humans have very few, scattered through the gland, with receptor expression that has been difficult to demonstrate convincingly in normal tissue. The FDA did the conservative thing in 2010, which was defensible when the human dataset was effectively zero patient-years. What is harder to defend is that the warning has not been revisited since.

The core claim The boxed warning describes a real experiment run in the wrong species. Sixteen years and several hundred thousand documented exposures later the human signal has not appeared, and the one signal that did appear turned out to be a picture of who got scanned rather than who got cancer.

The study that made everyone nervous

In 2023, Bezin and colleagues published a nested case-control analysis in Diabetes Care using the French national claims system, which captures essentially the entire population. They matched 2,562 incident thyroid cancers to 45,184 controls and applied a six-month lag to blunt protopathic bias.

The numbers were not subtle. Current GLP-1 RA use carried an adjusted hazard ratio of 1.46 (1.23 to 1.74) for all thyroid cancer, one to three years of cumulative use came in at 1.58 (1.27 to 1.95), and for medullary thyroid cancer, the histology the rodent work actually predicted, one to three years of use gave 1.78 (1.04 to 3.05).

That last figure is why the study mattered so much. A generic thyroid cancer signal is easy to wave away as surveillance. A medullary signal is mechanistically specific, and specificity is what makes a finding feel real. Several things undercut it, though, and they deserve to be named precisely rather than gestured at. The database carries no ICD code for medullary thyroid cancer, so the authors identified it by proxy through calcitonin testing and vandetanib prescriptions. Family history and radiation exposure were unmeasured. And the exposure window that produced the strongest signal, one to three years, is short for a solid tumor and long for a screening artifact, which is precisely the ambiguity that follow-up work had to resolve.

What three years of replication produced

It got resolved, more or less, and not in the direction the French study pointed.

Baxter and colleagues, Thyroid, 2025. Six national databases across Canada, Denmark, Norway, South Korea, Sweden, and Taiwan, using an active-comparator new-user design against DPP-4 inhibitors: roughly 98,000 GLP-1 RA users against about 2.5 million DPP-4i users, pooled weighted hazard ratio 0.81 (0.59 to 1.12). Sixty-seven thyroid cancer events across 362,436 person-years in the exposed group, which tells you both that the finding is null and that the study had no power for anything rare. Medullary subtype analysis was not possible at all.

One detail in that paper deserves more attention than it got. When the authors swapped the comparator from DPP-4 inhibitors to sulfonylureas, the estimate jumped to 1.80 (1.28 to 2.52). Same exposure, same outcome, different control group, opposite conclusion. The most plausible reading is confounding by obesity, since patients who get GLP-1s carry more adiposity than patients who get sulfonylureas, and obesity is itself an established risk factor for thyroid cancer. That single sensitivity analysis is a compact demonstration of why this literature has been so unstable.

Morales and colleagues, Diabetes Care, 2025. A federated analysis across international claims and EHR databases compared 460,032 GLP-1 RA users against 717,792 SGLT2 inhibitor users, 2,055,583 DPP-4i users, and 1,119,868 sulfonylurea users. Across all three pairwise comparisons the hazard ratios ran from 0.78 to 1.03, with intervals that comfortably included one, for thyroid tumors both benign and malignant.

Brito and colleagues, JAMA Otolaryngology, 2025. This is the study that explains the earlier signal rather than merely contradicting it. Using target trial emulation in a US claims cohort of more than 351,000 adults with type 2 diabetes, of whom 41,112 started a GLP-1 RA, the overall thyroid cancer hazard ratio was 1.24 (0.88 to 1.76). Null. But the year-by-year breakdown was the point: in the first twelve months, 1.85 (1.11 to 3.08), and after that, nothing.

Cancer does not behave that way. An agent that raises risk in month three and stops mattering by month fifteen is not a carcinogen, it is a magnifying glass, and the authors found the magnifying glass. Thyroid ultrasound was performed in 2.1% of GLP-1 starters by twelve months versus 1.5% of patients starting other agents. The boxed warning tells clinicians to worry about the thyroid, clinicians image the thyroid, imaging finds the subclinical papillary carcinomas that a large share of the adult population is quietly carrying, and the warning validates itself.

Annals of Internal Medicine, 2026. The randomized evidence, pooled: 48 placebo-controlled trials, 94,245 participants, median follow-up 70 weeks. Thyroid cancer occurred in 0.14% of GLP-1 RA recipients and 0.07% of placebo recipients, odds ratio 1.37 (0.82 to 2.31), graded moderate certainty for little or no effect. In absolute terms the interval spans roughly one fewer to nine additional cases per 10,000 patients.

Vilsbøll and colleagues, Diabetes, Obesity and Metabolism, February 2026. The largest integrated look yet, combining 93 clinical trials representing 101,732 participants and 206,950 patient-years with post-marketing surveillance and a US commercial claims database. Against placebo the hazard ratio was 1.70, with an interval running from 0.99 all the way out to 3.03, and the other trial comparisons landed in the same imprecise territory: 1.83 (0.70 to 6.71) against active comparators, and 1.41 (0.72 to 2.81) in the cardiovascular outcome trials. The real-world arm, compared against SGLT2 inhibitors, came back at 0.87 (0.58 to 1.29), while post-marketing surveillance reported thyroid cancer at 0.001 cases per 100 patient-years.

I want to be direct about that first number, because the paper concludes that the totality of the data does not suggest an association, and a lower bound of 0.99 is not the same thing as a lower bound of 0.60. It is a point estimate of 1.70 that misses conventional significance by the width of a rounding error, in an analysis whose author list includes Novo Nordisk affiliations. Read it as what it is: an imprecise estimate driven by a small number of events, compatible with no effect and also compatible with a modest one. The real-world arm of the same paper points the other way, which is itself informative about how much of this literature measures biology and how much measures who gets scanned.

The human evidence · thyroid cancer, GLP-1 RA vs comparator
One outlier, and everything since
Bezin 2023 · France, nested case-control
2,562 cases, 1–3 years of use
1.58 (1.27–1.95)
Baxter 2025 · six countries, vs DPP-4i
98,000 exposed, 362,436 person-years
0.81 (0.59–1.12)
Morales 2025 · three active comparators
460,032 exposed, 4.3M total
0.78–1.03
Brito 2025 · target trial emulation
year one, then everything after
1.85, then null
Ann Intern Med 2026 · 48 randomized trials
94,245 participants, median 70 weeks
1.37 (0.82–2.31)
Vilsbøll 2026 · 93 trials, vs placebo
206,950 patient-years
1.70 (0.99–3.03)
Comparators and designs differ across studies and the estimates are not pooled here. The two intervals that exclude 1.0 are the observational ones, and both come from designs that cannot separate a new cancer from a newly ordered ultrasound; the randomized pools sit above 1.0 on too few events to separate from chance.

The context that resolves most of it

Thyroid cancer is the textbook overdiagnosed malignancy. When South Korea added thyroid ultrasound to its national screening program, incidence rose roughly fifteenfold over two decades while mortality did not move at all, which Ahn, Kim, and Welch documented in the New England Journal in 2014. The tumors were always there. The screening found them.

Now layer that onto obesity medicine. Patients starting GLP-1s are heavier at baseline, and obesity independently raises thyroid cancer risk. They enter care more intensively than the general population, so they get more labs, more imaging, and more incidental findings, and they get all of it under a label that names thyroid tumors specifically. Every structural feature of the situation pushes toward finding more cancer in this group whether or not the drug does anything at all. Detection bias is not a hand-wave here. It is the most parsimonious explanation available, and unlike most explanations offered in this literature, it has been directly tested and supported.

Free for subscribers Every effect size in this piece, plus the ones from the trials that actually change prescribing, on one page you can keep next to your desk.

Get the GLP-1 Evidence Cheat Sheet →

The part that has not been answered

Medullary thyroid cancer remains genuinely untested in humans. It represents something like 1 to 2 percent of thyroid malignancies, which means that even two million patient-years produce too few events to analyze, and every study above says so explicitly. The honest position is that we have no human data on the outcome the rodent biology actually predicted, and observational sources may never deliver it.

Follow-up is also short. Median exposure in the international cohort ran 1.8 to 3.0 years, and the pooled randomized trials ran a median of 70 weeks, against a solid tumor latency measured in decades. Nearly all of this evidence comes from patients with type 2 diabetes, while the fastest-growing population on these drugs is younger, more often female, and does not have diabetes at all. That is the group whose data we most need and least have.

What would make me wrong
1. The outcome that matters has never been measured. Not one human cohort has had enough medullary thyroid cancer events to analyze. A null result for papillary cancer says nothing about the histology the rodent data actually predicted.
2. Both randomized pools sit above 1.0. 1.37 and 1.70, with intervals that would exclude the null on modestly more events. Underpowered is not the same as negative, and I would be reading these numbers very differently if the point estimates had landed at 0.9.
3. The follow-up and the population are both wrong. One to three years of exposure, overwhelmingly in type 2 diabetes, against a tumor whose latency runs in decades. Nobody has ten-year data in the non-diabetic obesity population now taking these drugs.

None of that makes the reassurance fake. It makes it partial. The claim the current evidence supports is narrow and worth stating exactly: in adults with type 2 diabetes followed for a few years, GLP-1 receptor agonists are not associated with more thyroid cancer than the drugs they compete with. Everything past that boundary is inference.

The orforglipron paradox

Here is where the regulatory logic stops making sense. Orforglipron, approved in 2026, is not pharmacologically active at the rat or mouse GLP-1 receptor. It produced no tumors in two-year rat carcinogenicity studies or in 26-week transgenic mouse studies, at exposures up to 26 times the maximum recommended human dose. There is no rodent C-cell finding to warn about, because the experiment that generated the entire class concern cannot be run on this molecule.

Its label carries the boxed warning regardless, and says so in as many words: the drug is not active in rats or mice and did not produce tumors, but the human relevance of GLP-1-receptor-dependent rodent C-cell tumors has not been determined, therefore the warning and the MEN2 contraindication apply. The January 2026 tirzepatide label carries the identical structure. At that point the boxed warning is no longer reporting a finding about a drug.

And it is not that the agency is incapable of revisiting these decisions. On January 13, 2026, the FDA requested removal of the suicidal ideation and behavior warning from liraglutide, semaglutide, and tirzepatide for weight management, citing a meta-analysis of 91 placebo-controlled trials in 107,910 patients alongside a claims cohort of 2.2 million. The machinery works. It simply has not been pointed at the thyroid.

Speculation · extrapolation beyond the data
If the warning is what drives the surveillance, then the warning has been quietly manufacturing its own evidence base for sixteen years, and every additional year it stands makes the observational literature harder to read rather than easier. The study that would settle the question is a medullary-specific registry with enough events to produce an interval narrower than a barn door. My guess is that nobody funds one, because the boxed warning already answers the question administratively and there is no regulatory pressure to replace an answer that costs the agency nothing. That is a hypothesis about institutional incentives, not a finding, and I would be glad to be wrong about it.

What I actually do

Here is what I actually do. I ask about personal and family history of medullary thyroid cancer and MEN2 at every GLP-1 consultation. A personal history of MTC or a known MEN2 syndrome ends the discussion: the contraindication is on the label, the alternative agents are adequate, and this is not the hill on which to test regulatory language. A family history starts a discussion rather than ending one. The label does not distinguish a germline RET syndrome from a single relative with apparently sporadic disease, but the biology does, since roughly 75 percent of MTC is sporadic and carries little heritable risk. So when the history looks sporadic, I say exactly that, and the patient and I decide together, weighing the label language against what the drug is likely to do for the problem in front of us. I do not order baseline calcitonin: the label itself calls the value of routine monitoring uncertain, calcitonin has a well documented false positive problem, and I see no reason to screen an unselected population for a cancer this rare. I do not order screening thyroid ultrasound, for the reason the Mayo group demonstrated: ordering it is the mechanism that manufactures the association we are trying to evaluate. When a nodule turns up incidentally, I work it up on standard criteria and I do not let the medication move the threshold, because deviating from that is how a null drug acquires a positive study.

And when a patient arrives having read that these drugs cause thyroid cancer, I tell them the version that is true, which is more interesting than either the scare or the reassurance. The warning is real and it comes from real experiments, those experiments were done in an animal whose thyroid is built differently from theirs, and sixteen years of looking for the human version has not turned it up. What we found instead is that if you tell doctors to look at a thyroid, they will find something. The accurate summary is not that these drugs are safe for the thyroid. It is that they have been examined hard, with steadily better methods, and nothing has emerged.

Taking a real family history before the first prescription takes ninety seconds and it is the only thyroid step in this article that actually changes a decision. Everything else on the label asks you to order tests that generate findings rather than answers. At Vineyard, that distinction is the job: our clinicians treat obesity as the chronic disease it is, and they spend the visit on the questions that change management.

See how Vineyard approaches obesity care →

Following the latest GLP-1 clinical trials?

Explore the trial tracker →

Disclosure: The author is Chief Medical Officer of Vineyard, a telehealth obesity medicine practice. This article is educational and is not individualized medical advice. Talk with your own clinician before making changes to your care.

REFERENCES

  1. Bjerre Knudsen L, Madsen LW, Andersen S, et al. Glucagon-like peptide-1 receptor agonists activate rodent thyroid C-cells causing calcitonin release and C-cell proliferation. Endocrinology. 2010;151(4):1473-1486. doi:10.1210/en.2009-1272

  2. Bezin J, Gouverneur A, Pénichon M, et al. GLP-1 receptor agonists and the risk of thyroid cancer. Diabetes Care. 2023;46(2):384-390. doi:10.2337/dc22-1148

  3. Ahn HS, Kim HJ, Welch HG. Korea's thyroid-cancer "epidemic": screening and overdiagnosis. N Engl J Med. 2014;371(19):1765-1767.

  4. Baxter SM, Lund LC, Andersen JH, et al. Glucagon-like peptide 1 receptor agonists and risk of thyroid cancer: an international multisite cohort study. Thyroid. 2025;35(1):69-78. doi:10.1089/thy.2024.0387

  5. Toro-Tobon D, Singh Ospina N, Brito JP. Thyroid cancer risk with GLP-1 receptor agonists: evidence, knowledge gaps, and the path forward. Thyroid. 2025;35(1):3-5. doi:10.1089/thy.2024.0690

  6. Brito JP, Herrin J, Swarna KS, et al. GLP-1RA use and thyroid cancer risk. JAMA Otolaryngol Head Neck Surg. 2025;151(3):243-252. doi:10.1001/jamaoto.2024.4852

  7. Morales DR, Bu F, Viernes B, et al. Risk of thyroid tumors with GLP-1 receptor agonists: a retrospective cohort study. Diabetes Care. 2025;48(8):1386-1394.

  8. Pollack R, Stokar J. Long-term glucagon-like peptide 1 receptor agonist use is not associated with increased risk of thyroid cancer in adults with type 2 diabetes. Diabetes Metab Res Rev. 2025;41(8):e70104. doi:10.1002/dmrr.70104

  9. Ko A, et al. Risk for cancer with glucagon-like peptide-1 receptor agonists and dual agonists: a systematic review and meta-analysis. Ann Intern Med. 2026;179(2):216-229. doi:10.7326/ANNALS-25-02237

  10. Vilsbøll T, Stellfeld M, Aroda VR, et al. Assessment of thyroid cancer risk associated with glucagon-like peptide 1 receptor agonist use. Diabetes Obes Metab. 2026;28(2):1499-1507. doi:10.1111/dom.70291

  11. Mannucci E, Dicembrini I. Glucagon-like peptide 1 receptor agonists and cancer risk: the good, the bad and the unknown. Nat Rev Clin Oncol. 2026;23(6):459-470.

  12. Eli Lilly and Company. MOUNJARO (tirzepatide) US Prescribing Information. Revised 01/2026.

  13. Eli Lilly and Company. FOUNDAYO (orforglipron) US Prescribing Information. Revised 07/2026.

  14. US Food and Drug Administration. FDA requests removal of suicidal behavior and ideation warning from glucagon-like peptide-1 receptor agonist (GLP-1 RA) medications. Drug Safety Communication. January 13, 2026.

Evidence-based obesity medicine, twice a week. No hype, no telehealth grifts.

Reply

Avatar

or to participate