Terence Tao on AI at ICM: Proofs will multiply, but mathematics won't accelerate
- AI is now solving research-grade mathematics. In First Proof's test of ten brand-new problems across four AI systems, seven yielded solutions at publishable quality. Yet on the Erdős problems website, nearly twenty AI-generated solutions sit unverified—no human expert is willing to review them. Solutions are multiplying, but downstream verification and digestion cannot keep up.
- At the International Congress of Mathematicians (ICM) public lecture on July 24, Terence Tao tackled this exact bottleneck. Bypassing debates over AI capability, he asked his audience to assume AI works—and then confront what mathematics must do next.
- Tao framed mathematical research as a five-step pipeline: Proof Generation → Verification → Exposition → Acceptance → Canonicalization. AI supercharges only the first step; every subsequent step remains bottlenecked by human judgment.
- Counterintuitively, overly smooth proofs can hinder learning. Tao displayed an annotated 1991 paper by Jean Bourgain from his youth, marked with "AARGH!" and "I hate Jean Bourgain." Those friction points were precisely what forced the reader to slow down and truly master the ideas.
- His prescription for the community comprises three rules: normalize AI usage disclosures, shift incentives from pure problem-solving to proof digestion, and enforce one golden rule: if you cannot clearly explain your result, you should not publish it.
AI is solving research-grade math. That's where the trouble begins
Dozens of AI-generated proofs currently pile up on the Erdős problems website.
Many of these proofs are likely correct. The catch is that nobody is verifying them.
Qualified mathematicians certainly exist, but few are willing to invest days wading through dozens of pages of machine-generated text to stake their reputation on a sign-off. In fact, several submitters explicitly noted in their submissions that they lacked the expertise to determine whether the proof was sound.
Addressing the audience, Tao posed a troubling question: Could we end up with a verified proof of a major landmark result that no human alive understands well enough to explain?
AI solving research-level mathematics is no longer a theoretical premise.
First Proof is an independent benchmark designed to measure AI performance on real research math. Mathematicians submit unsolved problems from their active work—problems solved privately but never published, making solutions impossible to find online or in the literature. In its second evaluation batch, ten new problems spanning computability theory, discrete geometry, stochastic PDEs, and von Neumann algebras were posed to four AI systems under strictly controlled conditions: one attempt per problem, no human intervention. Submissions underwent double-blind review by roughly thirty domain experts, with at least two reviewers per paper.
On a problem in stochastic PDEs, one AI system produced a solution entirely distinct from the human author's original approach, leaving reviewers thoroughly impressed.
On July 24 in Philadelphia, at the quadrennial International Congress of Mathematicians—the most prestigious gathering in the discipline—Tao delivered a public lecture titled "Mathematics in the Age of AI," accompanied by a 52-slide deck.
His core question: Now that AI can perform research mathematics, how must our discipline adapt?
The debate he refused to have
Tao structured the question of AI capability like a mathematical conjecture, intentionally leaving blanks throughout:
At some point in the near future, certain AI tools will, at some cost and under some level of human supervision, correctly perform certain research-grade mathematical tasks in certain fields, with some non-trivial success rate, accuracy, and quality.
He noted that splitting the conjecture into weak and strong versions suffices to show how differently mathematics must react:
Mathematicians can safely ignore AI as permanently irrelevant and carry on as usual.
Sustaining existing institutional culture becomes untenable—especially if our primary goal remains solving as many open problems as possible.
Tao then announced he would not debate the conjecture during his lecture.
His reasoning: the debate is at a stalemate. Data points on both sides abound, but few stem from controlled scientific trials. Public reports are distorted by reporting bias and non-scientific incentives, while key costs and variables remain unverified. He also cautioned against confusing wishful thinking with empirical reality: whether a statement is true is entirely separate from whether we want it to be true.
Instead, Tao asked the audience to adopt a "working hypothesis"—assuming AI capability as a given for conditional analysis. "I am not asking you to want it to be true, believe it to be true, or accept it as true," he clarified. "This is a conditional analysis. Evidence for or against the hypothesis is orthogonal to what I will discuss today."
A second crisis of foundations
The first foundational crisis struck a century ago. For centuries prior, mathematics relied on naive assumptions—leaving questions about sets, numbers, infinity, and axioms largely to philosophers. Russell's Paradox in 1901 and Gödel's Incompleteness Theorems in 1931 forced working mathematicians to re-examine their unstated premises.
Three decades of upheaval yielded an explicit, rigorous, and standardized axiomatic framework that withstood fierce scrutiny and serves as our trusted foundation today.
Tao argued that mathematics is entering a similar period of disruption. Where the previous crisis audited logical foundations, this one challenges institutional values and workflows. His conclusion was equally optimistic: once we rigorously re-evaluate and codify our practices, the mathematical community will emerge stronger and more resilient.
AI can solve more problems. But is that the real goal?
Rapid AI code generation is merely a symptom. Proofs bottleneck because solving problems has never been mathematics' sole objective.
Tao listed the motivations driving mathematical research—so numerous they filled an entire diagram:
Historically, these goals aligned: progress in one area naturally advanced the others. Consequently, surrogate metrics like problem count could serve as reliable proxies for broader health.
However, aggressive optimization triggers Goodhart's law: when a measure becomes a target, it ceases to be a good measure.
Tao added that generative AI is inherently ungrounded—lacking intrinsic truth-checking mechanisms—and when combined with commercial financial incentives, AI deployment becomes uniquely vulnerable to Goodhart's law.
Hyper-optimized AI tools cause previously aligned goals to diverge. Tao illustrated this shift across three sequential slides:
Goals were aligned. Advancing one naturally propelled the rest, making proxy metrics reliable.
Goals diverge in opposing directions. Proxy metrics fail: inflating solution counts no longer guarantees real progress elsewhere.
This dynamic extends far beyond mathematics. Every industry operates on aligned goals and reliance on proxy metrics. AI excels precisely at hyper-optimizing single metrics in isolation.
So what should our actual objectives be? In the core segment of his lecture, Tao redefined mathematical goals through five successive iterations.
From solving to textbooks: AI accelerates only step one
Tao focused on problem-solving—clarifying that while theory-building is equally vital, problem-solving is most immediately disrupted by AI.
Below is the five-stage pipeline resulting from his iterative refinements. Click through the tabs to watch it evolve:
v1: Solving Open Problems
This initial goal appears straightforward: open problems flow into solutions via proof generation.
The flaw was known long before AI: deluge of flawed solutions. Every number theorist regularly receives emails claiming to prove the Riemann Hypothesis.
v2: Adding Verification
The pipeline adds a second stage: generated outputs are merely unverified solutions until validated.
Here AI offers genuine assistance. Proof assistants like Lean, Rocq, and HOL verify proofs converted into line-by-line machine-checkable code. Using AI to automate this translation is known as autoformalization.
Requires an expert to spend days or weeks reading the proof, checking logic line by line, and risking their reputation on a sign-off.
Prone to fatigue and oversight, with little incentive for experts when reviewing uncredited work.
Translates proofs into languages like Lean for automated compiler checking, eliminating logical errors.
Requires autoformalization first—a computationally and labor-intensive process in its own right.
Tao observed that AI has significantly accelerated both generation and verification, a trend expected to continue under the working hypothesis.
He then posed the central dilemma: What happens when AI generates a lengthy proof that no human—including the person who prompted it—can understand?
This explains the backlog on the Erdős site: machine-verifiable proofs that lack any human willing to say, "I understand this and vouch for it."
v3: Adding Exposition
The pipeline expands again: verified solutions must be transformed into clear exposition.
In this phase, AI behavior becomes strikingly paradoxical.
Overly smooth proofs rob readers of understanding
Tao offered a sharp contrast when assessing AI's mathematical writing.
Spelling, grammar, and formatting are virtually flawless.
Tao added a footnote: arguably too flawless.
Laboriously expands on trivialities while glossing over—or obscuring—the most novel, intriguing steps of the argument.
Fails to contextualize findings within existing literature or provide high-level conceptual overviews.
Evaluators for First Proof noted the exact same pattern: AI systems obsess over routine calculations while hand-waving key breakthroughs with phrases like "follows by standard argument," without proof, or citing papers containing no such result.
A more troubling incident was documented: solutions borrowed verbatim phrasing, coined terminology (e.g., T-patterns, bends), and equation labels (B, T, D, H) from the author's previous paper without a single citation. Reviewers remarked that if submitted by a human, it would be flagged as outright plagiarism.
Hyper-optimized exposition
Exposition is a subjective target compared to formal verification. Even if AI improves its explanatory skills, hyper-optimized exposition presents a subtle hazard.
A proof can become excessively polished, flattening routine steps and profound insights into the same effortless tone.
In human writing, difficulty leaves authentic friction: convoluted syntax, skipped steps, strained phrasing. These cues warn the reader to slow down and scrutinize.
Excessive AI polishing erases both artificial friction (sloppy writing) and natural friction (inherent complexity). Readers glide through without grappling with core concepts. "Paradoxically," Tao noted, "the flaws in human exposition often serve the reader best."
Tao illustrated this point with a personal photograph.
The photograph serves as tangible proof of his thesis.
When Bourgain wrote "We skip the details," young Tao hit a wall. He underlined the text, added question marks, jotted "Sobolev norms!" to guide his intuition, and scribbled "I hate Jean Bourgain" in the margin.
Decades later, projecting that page at ICM, Tao underscored the point: those friction points forced him to pause, struggle, and truly master the mathematics. A perfectly uniform proof offers no such leverage.
He then quoted William Thurston's seminal 1994 essay, On Proof and Progress in Mathematics:
We are not taking orders for an abstract product composed of definitions, theorems and proofs. The measure of our success is whether what we do helps human beings understand and think about mathematics more clearly and effectively.
William Thurston, On Proof and Progress in Mathematics, 1994
Proofs Are Multiplying, but Mathematics Isn't Accelerating
Readable exposition is only step three.
v4: Community Acceptance
For a proof to advance mathematics, correctness and clarity are insufficient—other mathematicians must digest and integrate it into their own work.
Authors aid this process by sharing strategic insights: where they got stuck, why they chose specific pathways, and how key breakthroughs occurred. Proprietary AI models offer no such transparency into their internal reasoning.
Community acceptance is fundamentally slow and human. While lucid exposition accelerates it, acceptance remains a social process that cannot be unilaterally automated.
This ecosystem relies on volunteer efforts by human editors and peer reviewers. Though often viewed as less prestigious than generating original proofs, Tao emphasized that reviewing is indispensable for converting individual breakthroughs into collective progress.
AI can filter out inadequately verified or poorly written submissions, but filtering cannot replace community synthesis.
v5: Canonicalization
Publication is not the end of the line. Key results must eventually be codified into standard textbooks and foundational references—a process Tao called canonicalization.
As the slowest stage, canonicalization requires broad, deliberative consensus across the mathematical community, making it the least amenable to AI acceleration.
Yet Tao argued canonicalization is the most critical stage. First, practical applications become viable only after underlying mathematics is fully digested. Second, AI's present success relies directly on centuries of human canonicalization—models solve problems precisely because humans structured mathematics into accessible, coherent paradigms.
Diagnosis: Impedance Mismatch
Without institutional reform, uncalibrated AI integration creates an impedance mismatch—an electrical engineering concept describing energy loss when connecting incompatible circuits. In plain terms: severe proof indigestion.
The pipeline clogs at four critical junctions:
In short: we are transitioning from an era of proof scarcity to an era of proof glut.
These bottlenecks were accumulating long before modern AI, Tao noted.
The Culinary Analogy
Three months prior, Tao reframed the pipeline using cooking on social media:
In a food-scarce society, the bottleneck is foraging. While cleaning and cooking are appreciated, prestige belongs to hunters. Almost any edible game brought to the communal table is welcomed, and volunteers readily prepare it.
A potluck in an abundant society operates differently. Unsolicited raw game is unwelcome: a stranger dumping an uninspected carcass for others to clean earns no gratitude. Even pre-packaged, safety-inspected meals form only part of the table. Value resides in home-cooked dishes prepared by trusted community members—where conversations around the food nourish the community and train future chefs.
In that same thread, Tao offered an observation sharper than any slide: massive acceleration in proof generation has not meaningfully accelerated overall mathematical progress.
He also noted a perverse side effect: an AI "solution" that no human understands can kill interest in a problem—leaving it technically checked off while remaining completely uncomprehended.
What Mathematicians Must Do Next
With downstream bottlenecks clogging the pipeline, mathematicians must redirect their efforts.
Tao cited the Leiden Declaration as a constructive starting point. Released on June 2, 2026, and drafted by sixteen mathematicians from Cambridge, Columbia, Oxford, Leiden, ETH Zürich, and other institutions, the declaration has been endorsed by the International Mathematical Union (IMU).
Tao highlighted four operational principles from the declaration, paired with his own commentary:
The full Leiden Declaration encompasses broader issues—including AI in warfare, surveillance, democratic disruption, and environmental impact—calling for public oversight under the advice "Don't believe the hype." Tao focused strictly on operational guidelines for working mathematicians.
Practicing What He Preaches
Tao included two personal disclosures in his slides. Slide 1 footnoted: "All em-dashes in these slides were typed by hand." (Long em-dashes are often cited as telltale markers of AI prose.)
On the slide advocating disclosure normalization, his footnote read: "These slides used AI tools for text autocompletion and diagram generation."
Both notes were placed unobtrusively in the footers.
Conclusion
Tao concluded by stressing that problem-solving is only one aspect of mathematics. Similar analyses must be conducted for teaching, mentoring, hiring, grant allocation, and public outreach. In domains like education, human element must be safeguarded with strict limits on AI; in others, mathematicians must proactively define AI integration on their own terms. Developing new infrastructure to complement traditional workflows remains an essential task for future lectures.
His closing remark to the hall: "Our community needs to sit down together and engage in open, honest discussion regarding AI capabilities alongside our shared goals and values."
For his final slide, Tao added no further commentary, presenting only Article 7 of the Leiden Declaration:
AI Solves Research-Grade Math. The Bottleneck Is the Next Four Steps
Terence Tao deconstructs mathematical research into a five-step pipeline at ICM. Here is the full summary in a single visual page.
↓ One-page summary · Includes an animated diagram
Dozens of AI-generated proofs sit on the Erdős problems website (a repository collecting open prize problems posed by Paul Erdős). Many may be valid, but lack human experts willing to read them, confirm validity, and sign off. Several submitters explicitly noted they lacked qualifications to evaluate their own entries.
AI capability in research math is already reality. First Proof, an independent benchmark, collects unreleased problems directly from active research, presenting them to AI systems under controlled trial conditions.
Data sourced from First Proof Report #2 and the Erdős problems site. In his July 24 ICM lecture, Tao cited headline numbers, bypassing capability debates to ask: Assuming AI works, how must mathematics adapt?
Tao cataloged core mathematical pursuits: solving open problems, building theory, understanding the world, cultivating community, mentoring students, and crafting enduring aesthetic works. Historically, these goals aligned, allowing solution counts to serve as a reliable proxy.
Goodhart's law warns that when a measure becomes a target, it ceases to be a good measure. Unanchored generative AI outputs combined with commercial incentives make math research uniquely susceptible to this distortion.
Starting from "solving open problems," Tao added a requirement at each flaw: verify correctness, ensure human readability, achieve peer acceptance, and integrate into canonical theory (standard textbook material). Through five iterations, a five-step pipeline emerged—shrinking AI's autonomous share at every stage.
Tao noted AI writing is split: pristine grammar and formatting ("arguably too flawless"), yet verbose on trivialities while obscuring essential breakthroughs. First Proof reviewers observed the same: key steps dismissed as "standard arguments," and uncredited verbatim borrowings from authors' prior papers.
Human proofs retain natural friction where authors struggled—stiff phrasing and concise skips signal readers to pause. Excessive AI polishing flattens routine steps and deep insights into uniform smoothness.
Tao endorsed the Leiden Declaration (issued June 2, 2026; backed by IMU), extracting four practical guidelines: disclose tool usage, facilitate peer review, retain human accountability for correctness, and ensure proper attribution.
Race to claim priority in solving open problems.
Exposition, review, and canonicalization are seen as less prestigious, relying on uncredited volunteer labor.
De-emphasize raw generation and priority claims; reward proof digestion—exposition, review, and canonicalization.
Normalize responsible disclosures to prevent covert AI usage hidden to avoid peer criticism.
Rule of thumb: If you cannot clearly explain your result, do not publish it.
Tao led by example with footer disclosures: manual em-dashes and AI assistance for text and diagrams. Similar audits are needed across teaching, mentoring, hiring, grants, and public communication.
zero sign-offs!
At least one AI
reached publishable
quality on seven.
is math stuck?
whether AI works.
do next?
then ask the question.
built the full
pipeline live
on stage.
is a drawback?!
understanding.
who will digest?
- × Blocked before verification
- × Blocked before exposition
- × Blocked before review
- × Blocked before canonicalization
doesn't speed up math.
mathematicians do?
Do not publish.
where humans govern.