Deep Dive · Xiaohu Analysis

Terence Tao on AI at ICM: Proofs will multiply, but mathematics won't accelerate

Breaking research into a five-step pipeline, Terence Tao shows AI accelerates only the first step. The remaining four grow progressively slower and more human-dependent.
Executive Summary
  • AI is now solving research-grade mathematics. In First Proof's test of ten brand-new problems across four AI systems, seven yielded solutions at publishable quality. Yet on the Erdős problems website, nearly twenty AI-generated solutions sit unverified—no human expert is willing to review them. Solutions are multiplying, but downstream verification and digestion cannot keep up.
  • At the International Congress of Mathematicians (ICM) public lecture on July 24, Terence Tao tackled this exact bottleneck. Bypassing debates over AI capability, he asked his audience to assume AI works—and then confront what mathematics must do next.
  • Tao framed mathematical research as a five-step pipeline: Proof Generation → Verification → Exposition → Acceptance → Canonicalization. AI supercharges only the first step; every subsequent step remains bottlenecked by human judgment.
  • Counterintuitively, overly smooth proofs can hinder learning. Tao displayed an annotated 1991 paper by Jean Bourgain from his youth, marked with "AARGH!" and "I hate Jean Bourgain." Those friction points were precisely what forced the reader to slow down and truly master the ideas.
  • His prescription for the community comprises three rules: normalize AI usage disclosures, shift incentives from pure problem-solving to proof digestion, and enforce one golden rule: if you cannot clearly explain your result, you should not publish it.
The Scene

AI is solving research-grade math. That's where the trouble begins

Dozens of AI-generated proofs currently pile up on the Erdős problems website.

Many of these proofs are likely correct. The catch is that nobody is verifying them.

Qualified mathematicians certainly exist, but few are willing to invest days wading through dozens of pages of machine-generated text to stake their reputation on a sign-off. In fact, several submitters explicitly noted in their submissions that they lacked the expertise to determine whether the proof was sound.

Addressing the audience, Tao posed a troubling question: Could we end up with a verified proof of a major landmark result that no human alive understands well enough to explain?

AI solving research-level mathematics is no longer a theoretical premise.

First Proof is an independent benchmark designed to measure AI performance on real research math. Mathematicians submit unsolved problems from their active work—problems solved privately but never published, making solutions impossible to find online or in the literature. In its second evaluation batch, ten new problems spanning computability theory, discrete geometry, stochastic PDEs, and von Neumann algebras were posed to four AI systems under strictly controlled conditions: one attempt per problem, no human intervention. Submissions underwent double-blind review by roughly thirty domain experts, with at least two reviewers per paper.

7 / 10
Problems where at least one AI system produced publishable-quality solutions
39
Total solution papers submitted across four systems, double-blind reviewed by ~30 experts
1
Problem where all four systems failed to make meaningful progress (metric geometry)

On a problem in stochastic PDEs, one AI system produced a solution entirely distinct from the human author's original approach, leaving reviewers thoroughly impressed.

On July 24 in Philadelphia, at the quadrennial International Congress of Mathematicians—the most prestigious gathering in the discipline—Tao delivered a public lecture titled "Mathematics in the Age of AI," accompanied by a 52-slide deck.

His core question: Now that AI can perform research mathematics, how must our discipline adapt?

The debate he refused to have

Tao structured the question of AI capability like a mathematical conjecture, intentionally leaving blanks throughout:

AI Capability Conjecture · Template

At some point in the near future, certain AI tools will, at some cost and under some level of human supervision, correctly perform certain research-grade mathematical tasks in certain fields, with some non-trivial success rate, accuracy, and quality.

The highlighted spans are placeholders. How you fill them defines your specific version of the conjecture. Debate over AI capability often stems from people filling different blanks.

He noted that splitting the conjecture into weak and strong versions suffices to show how differently mathematics must react:

If even the weak form fails

Mathematicians can safely ignore AI as permanently irrelevant and carry on as usual.

If the strongest form holds

Sustaining existing institutional culture becomes untenable—especially if our primary goal remains solving as many open problems as possible.

Tao then announced he would not debate the conjecture during his lecture.

His reasoning: the debate is at a stalemate. Data points on both sides abound, but few stem from controlled scientific trials. Public reports are distorted by reporting bias and non-scientific incentives, while key costs and variables remain unverified. He also cautioned against confusing wishful thinking with empirical reality: whether a statement is true is entirely separate from whether we want it to be true.

Instead, Tao asked the audience to adopt a "working hypothesis"—assuming AI capability as a given for conditional analysis. "I am not asking you to want it to be true, believe it to be true, or accept it as true," he clarified. "This is a conditional analysis. Evidence for or against the hypothesis is orthogonal to what I will discuss today."

A second crisis of foundations

The first foundational crisis struck a century ago. For centuries prior, mathematics relied on naive assumptions—leaving questions about sets, numbers, infinity, and axioms largely to philosophers. Russell's Paradox in 1901 and Gödel's Incompleteness Theorems in 1931 forced working mathematicians to re-examine their unstated premises.

Three decades of upheaval yielded an explicit, rigorous, and standardized axiomatic framework that withstood fierce scrutiny and serves as our trusted foundation today.

Tao argued that mathematics is entering a similar period of disruption. Where the previous crisis audited logical foundations, this one challenges institutional values and workflows. His conclusion was equally optimistic: once we rigorously re-evaluate and codify our practices, the mathematical community will emerge stronger and more resilient.

Mechanism

AI can solve more problems. But is that the real goal?

Rapid AI code generation is merely a symptom. Proofs bottleneck because solving problems has never been mathematics' sole objective.

Tao listed the motivations driving mathematical research—so numerous they filled an entire diagram:

Build Theories
Solve Olympiad Problems
Solve Erdős Problems
Apply Knowledge
Mathematical Knowledge↖ ↑ ↗ ← → ↙ ↓ ↘
Solve Research Problems
Build Community
Train Next Generation
Create Enduring Aesthetic Value
Redrawn based on Slide 22, listing Tao's core motivations for mathematics. The eight radiating arrows from the center represent alignment across goals. Tao noted in a footnote that this diagram is highly simplified from a much higher-dimensional reality.

Historically, these goals aligned: progress in one area naturally advanced the others. Consequently, surrogate metrics like problem count could serve as reliable proxies for broader health.

However, aggressive optimization triggers Goodhart's law: when a measure becomes a target, it ceases to be a good measure.

Tao added that generative AI is inherently ungrounded—lacking intrinsic truth-checking mechanisms—and when combined with commercial financial incentives, AI deployment becomes uniquely vulnerable to Goodhart's law.

Hyper-optimized AI tools cause previously aligned goals to diverge. Tao illustrated this shift across three sequential slides:

Historically

Goals were aligned. Advancing one naturally propelled the rest, making proxy metrics reliable.

Under hyper-optimization

Goals diverge in opposing directions. Proxy metrics fail: inflating solution counts no longer guarantees real progress elsewhere.

Redrawn from Slides 20–22 showing goal divergence. Both central hubs represent "Mathematical Knowledge."

This dynamic extends far beyond mathematics. Every industry operates on aligned goals and reliance on proxy metrics. AI excels precisely at hyper-optimizing single metrics in isolation.

So what should our actual objectives be? In the core segment of his lecture, Tao redefined mathematical goals through five successive iterations.

Deconstruction

From solving to textbooks: AI accelerates only step one

Tao focused on problem-solving—clarifying that while theory-building is equally vital, problem-solving is most immediately disrupted by AI.

Below is the five-stage pipeline resulting from his iterative refinements. Click through the tabs to watch it evolve:

Goal · v1Solve as many open problems as possible.
Goal · v2Solve as many open problems as possible and verify that the solutions are correct.
Goal · v3Solve open problems, verify correctness, and ensure results are clearly communicated and understood by the mathematical community.
Goal · v4Solve open problems, verify correctness, communicate clearly, and have results digested and accepted by the community.
Goal · v5Solve open problems, verify correctness, communicate clearly, digest and accept results, and integrate them into canonical theory.
Open Problems
Proof GenerationDramatically accelerated by AI
Solutions
Proof VerificationAI helps; formal proof assistants verify
Verified Solutions
Proof ExpositionAI track record is mixed
Expository Solutions
Proof PublicationSlow; relies on human editors & reviewers
Accepted Solutions
Proof CanonicalizationSlowest; requires deliberative community consensus
Canonical Theory
Combined interactive diagram based on Slides 24, 26, 29, 37, and 43. Annotations on AI suitability synthesize Tao's commentary across these slides.

v1: Solving Open Problems

This initial goal appears straightforward: open problems flow into solutions via proof generation.

The flaw was known long before AI: deluge of flawed solutions. Every number theorist regularly receives emails claiming to prove the Riemann Hypothesis.

v2: Adding Verification

The pipeline adds a second stage: generated outputs are merely unverified solutions until validated.

Here AI offers genuine assistance. Proof assistants like Lean, Rocq, and HOL verify proofs converted into line-by-line machine-checkable code. Using AI to automate this translation is known as autoformalization.

Human Verification

Requires an expert to spend days or weeks reading the proof, checking logic line by line, and risking their reputation on a sign-off.

Prone to fatigue and oversight, with little incentive for experts when reviewing uncredited work.

Machine Verification

Translates proofs into languages like Lean for automated compiler checking, eliminating logical errors.

Requires autoformalization first—a computationally and labor-intensive process in its own right.

Tao observed that AI has significantly accelerated both generation and verification, a trend expected to continue under the working hypothesis.

He then posed the central dilemma: What happens when AI generates a lengthy proof that no human—including the person who prompted it—can understand?

This explains the backlog on the Erdős site: machine-verifiable proofs that lack any human willing to say, "I understand this and vouch for it."

v3: Adding Exposition

The pipeline expands again: verified solutions must be transformed into clear exposition.

In this phase, AI behavior becomes strikingly paradoxical.

Counterintuitive

Overly smooth proofs rob readers of understanding

Tao offered a sharp contrast when assessing AI's mathematical writing.

The Good

Spelling, grammar, and formatting are virtually flawless.

Tao added a footnote: arguably too flawless.

The Bad

Laboriously expands on trivialities while glossing over—or obscuring—the most novel, intriguing steps of the argument.

Fails to contextualize findings within existing literature or provide high-level conceptual overviews.

Evaluators for First Proof noted the exact same pattern: AI systems obsess over routine calculations while hand-waving key breakthroughs with phrases like "follows by standard argument," without proof, or citing papers containing no such result.

A more troubling incident was documented: solutions borrowed verbatim phrasing, coined terminology (e.g., T-patterns, bends), and equation labels (B, T, D, H) from the author's previous paper without a single citation. Reviewers remarked that if submitted by a human, it would be flagged as outright plagiarism.

Hyper-optimized exposition

Exposition is a subjective target compared to formal verification. Even if AI improves its explanatory skills, hyper-optimized exposition presents a subtle hazard.

A proof can become excessively polished, flattening routine steps and profound insights into the same effortless tone.

In human writing, difficulty leaves authentic friction: convoluted syntax, skipped steps, strained phrasing. These cues warn the reader to slow down and scrutinize.

Excessive AI polishing erases both artificial friction (sloppy writing) and natural friction (inherent complexity). Readers glide through without grappling with core concepts. "Paradoxically," Tao noted, "the flaws in human exposition often serve the reader best."

Human-written proof
Curve height indicates reading fluency. Dips represent spots where the author struggled—stiff phrasing and concise leaps force readers to pause and digest core ideas.
Hyper-polished AI proof
Uniformly smooth from start to finish. Readers glide through without internalizing key concepts.
Conceptual diagram created by Xiaohu Analysis (not present in slide deck).

Tao illustrated this point with a personal photograph.

A spread of a 1991 paper by Jean Bourgain, covered with dense pen annotations in the margins and between lines
Slide 33 caption: "A 1991 paper by Jean Bourgain, heavily annotated by a much younger version of myself." Jean Bourgain (1954–2018) was a Fields Medalist and giant in harmonic analysis.
Close-up of the same page: printed phrase 'We skip the details' underlined with handwritten 'AARGH!', three question marks, and 'I hate Jean Bourgain'
Close-up excerpt. Underneath the printed line "We skip the details.", young Tao underlined the text and wrote AARGH!, three question marks, "Sobolev norms!", and "I hate Jean Bourgain."

The photograph serves as tangible proof of his thesis.

When Bourgain wrote "We skip the details," young Tao hit a wall. He underlined the text, added question marks, jotted "Sobolev norms!" to guide his intuition, and scribbled "I hate Jean Bourgain" in the margin.

Decades later, projecting that page at ICM, Tao underscored the point: those friction points forced him to pause, struggle, and truly master the mathematics. A perfectly uniform proof offers no such leverage.

He then quoted William Thurston's seminal 1994 essay, On Proof and Progress in Mathematics:

We are not taking orders for an abstract product composed of definitions, theorems and proofs. The measure of our success is whether what we do helps human beings understand and think about mathematics more clearly and effectively.

William Thurston, On Proof and Progress in Mathematics, 1994
Diagnosis

Proofs Are Multiplying, but Mathematics Isn't Accelerating

Readable exposition is only step three.

v4: Community Acceptance

For a proof to advance mathematics, correctness and clarity are insufficient—other mathematicians must digest and integrate it into their own work.

Authors aid this process by sharing strategic insights: where they got stuck, why they chose specific pathways, and how key breakthroughs occurred. Proprietary AI models offer no such transparency into their internal reasoning.

Community acceptance is fundamentally slow and human. While lucid exposition accelerates it, acceptance remains a social process that cannot be unilaterally automated.

This ecosystem relies on volunteer efforts by human editors and peer reviewers. Though often viewed as less prestigious than generating original proofs, Tao emphasized that reviewing is indispensable for converting individual breakthroughs into collective progress.

AI can filter out inadequately verified or poorly written submissions, but filtering cannot replace community synthesis.

v5: Canonicalization

Publication is not the end of the line. Key results must eventually be codified into standard textbooks and foundational references—a process Tao called canonicalization.

As the slowest stage, canonicalization requires broad, deliberative consensus across the mathematical community, making it the least amenable to AI acceleration.

Yet Tao argued canonicalization is the most critical stage. First, practical applications become viable only after underlying mathematics is fully digested. Second, AI's present success relies directly on centuries of human canonicalization—models solve problems precisely because humans structured mathematics into accessible, coherent paradigms.

Diagnosis: Impedance Mismatch

Without institutional reform, uncalibrated AI integration creates an impedance mismatch—an electrical engineering concept describing energy loss when connecting incompatible circuits. In plain terms: severe proof indigestion.

The pipeline clogs at four critical junctions:

Before VerificationAI-generated proofs pile up awaiting human validation.
Before ExpositionVerified proofs languish without readable human summaries.
Before Peer ReviewEven correct and well-written submissions overwhelm traditional peer review.
Before CanonicalizationPublished results proliferate faster than the community can synthesize them into core theory.

In short: we are transitioning from an era of proof scarcity to an era of proof glut.

Proof GenerationDramatically accelerated by AI
Proof VerificationFormal tools help, but humans prioritize what's worth verifying
Proof DigestionExposition, review, canonicalization proceed at human speed
Diagram of velocity differential across pipeline stages (created by Xiaohu Analysis). Bar lengths indicate relative speed; hatched areas mark backlogs at junctions.

These bottlenecks were accumulating long before modern AI, Tao noted.

The Culinary Analogy

Three months prior, Tao reframed the pipeline using cooking on social media:

Proof Generation = Foraging Proof Verification = Cleaning & Inspection Proof Digestion = Cooking

In a food-scarce society, the bottleneck is foraging. While cleaning and cooking are appreciated, prestige belongs to hunters. Almost any edible game brought to the communal table is welcomed, and volunteers readily prepare it.

A potluck in an abundant society operates differently. Unsolicited raw game is unwelcome: a stranger dumping an uninspected carcass for others to clean earns no gratitude. Even pre-packaged, safety-inspected meals form only part of the table. Value resides in home-cooked dishes prepared by trusted community members—where conversations around the food nourish the community and train future chefs.

Excerpted from Terence Tao's Mastodon post (April 27, 2026). Included here for conceptual clarity.

In that same thread, Tao offered an observation sharper than any slide: massive acceleration in proof generation has not meaningfully accelerated overall mathematical progress.

He also noted a perverse side effect: an AI "solution" that no human understands can kill interest in a problem—leaving it technically checked off while remaining completely uncomprehended.

Recommendations

What Mathematicians Must Do Next

With downstream bottlenecks clogging the pipeline, mathematicians must redirect their efforts.

Tao cited the Leiden Declaration as a constructive starting point. Released on June 2, 2026, and drafted by sixteen mathematicians from Cambridge, Columbia, Oxford, Leiden, ETH Zürich, and other institutions, the declaration has been endorsed by the International Mathematical Union (IMU).

Tao highlighted four operational principles from the declaration, paired with his own commentary:

Declaration · Article 1Disclose Tool Usage
Transparently disclose the use of automated tools, including LLMs, machine learning systems, proof assistants, and software. Include a "Tools and Computational Resources" section in papers. When reviewing, adhere to publisher policies; if AI is permitted, disclose usage and remain personally accountable for your evaluation.
Terence Tao's CommentaryWe must prevent the worst-case scenario: authors covertly using AI while hiding it to avoid peer criticism. Responsible disclosure must be normalized.
Declaration · Article 2Support Reviewer Needs
Using AI in paper preparation can add friction to peer review. Facilitate review by disclosing tool usage, providing complete citations of prior work, and submitting formal proofs whenever feasible.
Terence Tao's CommentaryWe must de-emphasize raw proof generation and claims of priority, shifting focus to proof digestion—exposition, review, and canonicalization.
Declaration · Articles 4 & 6Retain Responsibility for Correctness · Ensure Proper Attribution
Human authors retain full accountability for the validity of arguments, accuracy of results, and completeness of citations when employing automated tools. Known limitations of AI in tracking intellectual lineage create an affirmative duty to actively trace and credit foundational sources.
Terence Tao's Rule of ThumbIf authors cannot convincingly demonstrate the ability to deliver a clear, expert-level, accurate, and properly attributed presentation of their result, that result should not be published.

The full Leiden Declaration encompasses broader issues—including AI in warfare, surveillance, democratic disruption, and environmental impact—calling for public oversight under the advice "Don't believe the hype." Tao focused strictly on operational guidelines for working mathematicians.

Practicing What He Preaches

Tao included two personal disclosures in his slides. Slide 1 footnoted: "All em-dashes in these slides were typed by hand." (Long em-dashes are often cited as telltale markers of AI prose.)

On the slide advocating disclosure normalization, his footnote read: "These slides used AI tools for text autocompletion and diagram generation."

Both notes were placed unobtrusively in the footers.

Conclusion

Tao concluded by stressing that problem-solving is only one aspect of mathematics. Similar analyses must be conducted for teaching, mentoring, hiring, grant allocation, and public outreach. In domains like education, human element must be safeguarded with strict limits on AI; in others, mathematicians must proactively define AI integration on their own terms. Developing new infrastructure to complement traditional workflows remains an essential task for future lectures.

His closing remark to the hall: "Our community needs to sit down together and engage in open, honest discussion regarding AI capabilities alongside our shared goals and values."

For his final slide, Tao added no further commentary, presenting only Article 7 of the Leiden Declaration:

Screenshot of original text of Leiden Declaration Article 7: Participate in public discourse
Slide 51, the final slide of the lecture, displaying only Article 7 without added text. English text below.
Declaration · Article 7Participate in Public Discourse
Mathematicians have a duty to support rigorous science journalism and participate in public discourse to contextualize AI-assisted methods and results. This is vital within specialized subfields where assessing depth, difficulty, and significance requires domain expertise. Mathematicians are encouraged to collaborate with and support researchers facing similar challenges.
The final slide also listed notable emerging infrastructures: Mathlib (Lean's mathematical library), the Erdős Problems website, Tao's crowdsourced database of mathematical constants, and open competitions by the SAIR Foundation (co-founded by Tao). An entry titled "Mathematical Discourse" was marked TBA.
Source
Mathematics in the age of AITerence Tao · ICM 2026 Public Lecture·Slide Deck PDF·July 24, 2026
Publisher's Note
The two annotated paper photographs stem from Slide 33 (the second is a cropped close-up by Xiaohu Analysis). The five-step pipeline and goal divergence diagrams are English re-renderings based on Slides 22, 24, 26, 29, 37, and 43. The velocity differential chart is a conceptual illustration and does not represent empirical measurement. First Proof review details, domain areas, and ratings stem from its second report (arXiv:2606.18119); Tao's slides cited only headline conclusions. Per-problem costs ranged from $8 to $951 in the report ($10–$1,000 cited in lecture). Note that Terence Tao is a team member for UCLA Moonshot Harness (one of the four evaluated systems), though this was unmentioned in his slides. The culinary analogy originated in his April 27, 2026 Mastodon post, not the ICM lecture.