The potential impact of AI on mathematics has received much attention recently. Such conversations occur in many different settings, which naturally leads to a somewhat fragmented scattering of opinions. This makes it all the more relevant that a cohesive and balanced perspective was recently composed in the Leiden Declaration on Artificial Intelligence and Mathematics.
A key challenge naturally remains to implement this all concretely in our community norms and institutional policies. Having a document with clear support across the community can only help with this, as it gives something explicit that we can all agree with. I would hence like to encourage any reader to consult the declaration and, if they agree, sign it at the following link: https://leidendeclaration.ai/
One of the recommendations in the declarations is for individual mathematicians to participate in the public discourse. This post is my contribution in this regard.
Why and how should we do mathematics?
Before one can discuss what consequences of a technological development are positive or negative, one has to define what is considered desirable. The declaration hence starts with a section on our values.
I will categorize these values in two broad categories:
- Why we do mathematics. These motivations includes curiosity as well as concrete applications. Crucially for these motivations, mathematics is not just a sequence of dry formal statements and proofs. We desire to gain understanding into the phenomena being studied. Further, the selection of new directions requires us to develop judgement into what are fruitful questions.
- Responsibilities that have to be fulfilled for the persistence of a community that can achieve the motivations of the first bullet. These include giving verifiable proofs and correct attribution of ideas to previous literature, as well as accurately evaluating others’ work when refereeing. I would also add nurturing future generations, which includes good teaching and mentoring.
I think that it is here correct that the declaration also emphasize the non-formal aspects of our motivations. Understanding itself is indeed a crucial output of both applied and pure research. In what follows, I would like to discuss this point in a bit more detail.
In an applied setting, the fact that one should not lose track of understanding was also stressed in the context of a previous significant technological development. In the opening paragraph his 1962 book on numerical methods, Richard Hamming stressed that the purpose of computing is insight, not numbers. The full quote goes as follows:
Numerical methods use numbers to simulate mathematical processes, which in turn usually simulate real-world situations. This implies that there is a purpose behind the computing. To cite the motto of the book, The Purpose of Computing is Insight, Not Numbers. This motto is often thought to mean that the numbers from a computing machine should be read and used, but there is much more to the motto. The choice of the particular formula, or algorithm, influences not only the computing but also how we are to understand the results when they are obtained…
– Richard Hamming (1962)
Note that this book was published more than 60 years ago. It is striking how relevant this quote remains in our present setting!
Computing can be a useful tool to explore various phenomena, but this does not make humans redundant. It is the interpretation of the results that is valuable. Moreover, the challenges of algorithmic implementation can lead to new insight. For a simple illustrative example in a strictly mathematical setting, recall that one way to define the coprimality of two integers $a,b \in \mathbb{Z}$ is to state that their the greatest common divisor equals one. An equivalent definition would be that they do not share prime divisors. However, while computing the greatest common divisor is accomplished by the Euclidean algorithm, prime factorization is typically difficult. An appreciation for different ways to define a notion thus arise naturally from challenges that one may encounter if one tries to implement algorithms.
More generally, the challenges involved in solving a mathematical problem frequently lead to new insight. Moreover, sometimes it is exactly the suffering with the tedious parts of a problem that leads us to a crucial innovation. For instance, one may recognize that one is studying the “wrong problem”. By considering a different problem that shares features with the original problem but is simpler in certain respects, it may become possible to get a clearer picture of the phenomenon that we actually wanted to understand. A brute force solution to the original problem would then lead to less insight. One potential downside of over-reliance on AI is that too smooth an experience may reduce the opportunity that such challenges offer.
On the other hand, I would certainly not claim that friction is always helpful. Some directions may remain unexplored due to many but unavoidable small frictions that do not themselves yield new insight, but stand in the way of the key challenge that does. In other cases, it can occur that the tools required to understand the key challenge itself are intimidating. A more optimistic view it that appropriate use of AI could potentially help overcome such barriers and thus allow us to penetrate closer to the center of the matter.
Another perspective on the crucial value of understanding can be found in the wonderful essay On proofs and progress in mathematics by William Thurston from 1994, who defined the work of a mathematicians circularly as the advancement of the human understanding of mathematics.
[…] It would not be good to start, for example, with the question
How do mathematicians prove theorems?
This question introduces an interesting topic, but to start with it would be to project two hidden assumptions:
- that there is uniform, objective and firmly established theory and practice of mathematical proof, and
- that progress made by mathematicians consists of proving theorems.
It is worthwhile to examine these hypotheses, rather than to accept them as obvious and proceed from there. […], as a more explicit (and leading) form of the question, I prefer
How do mathematicians advance human understanding of mathe-
matics?
This question brings to the fore something that is fundamental and pervasive: that what we are doing is finding ways for people to understand and think about mathematics.– William Thurston (1994)
This quote also occurred in the context of a debate about the standards of the mathematical community. It was an invited response to a 1993 paper by Jaffe and Quinn titled “Theoretical mathematics”: towards a cultural synthesis of mathematics and theoretical phyisics. The latter paper cautioned about the dangers of works on the connection between mathematics and physics that were not completely rigorous, and advocated for a clean separation between rigorous work and mere speculation.
Thurston’s response emphasizes that the projection on a one-dimensional scale of speculation versus rigor ignores many basic phenomena. For example, the following aspects are also relevant for our current discussion:
- Concerning proofs, what is most important to us is that they are understandable by humans. This is quite different from completely formal proofs that can be verified from a machine. Making a proof understandable for humans requires communicating the few crucial ideas and the global proof architecture. Formal proofs, on the other hands, can necessitate tremendous bookkeeping of tiny details that may distract from the key point.
- The strong emphasis on “theorem credits” in how we evaluate mathematicians has a negative effect. Mathematics occurs in a community and there are ecological effects. For instance, a sequence of impressive results can prematurely kill a field because other people feel that it will be difficult to gain theorem credits, or because the underlying ideas of the arguments are not transparent to others.
The distinction between human and formal proofs has direct consequences for AI systems. Formalizing the output of large language models in proof assistants like Lean would be useful to reduce the amount of noise from false proofs, and to reduce the pressures on the overloaded refereeing system. However, a formal proof in isolation would not be sufficient to lead to human understanding. The main value of the argument and result come from its digestion by humans. By definition, the latter requires human involvement.
A correct proof generated by an autonomous AI system may sometimes be worse than no proof at all, since it removes an incentive for humans to put effort into understanding the topic. A related issue applies when someone announces an AI-generated proof with insufficient effort in the verification and digestion of the arguments. There is an externality: the community is burdened by a task that should properly be done by the one claiming credit. The credit-based system makes it is difficult for others to justify spending the effort, even if this is crucial for the health of the community. These negative aspects of the theorem economy predate AI, but they are further amplified by it.
Recommendations
The main substance of the Leiden declaration may be the recommendations that it outlines. There are four large categories depending on the target group: individual mathematicians, mathematical institutions and funders, governments, and AI companies.
Some of the recommendations are directly implied by the responsibilities outlined in the previous section, with the main issue being that inappropriate use of AI systems may lead to neglect of these responsibilities and our community norms. One that is not yet common practice and hence strikes me as particularly useful and actionable is the following recommendation for authors:
Disclose tool use
Transparently disclose the use of automated tools, including large language models, machine learning systems, proof assistants, and other mathematical software. Include a “Tool and computational resource disclosure” section in your papers…
– Leiden Declaration on Artificial Intelligence and Mathematics (2026)
The declaration goes on by acknowledging that the precise form of such a section will naturally evolve, and recommends consulting journal policies as well as the spirit of the UNESCO Recommendation on Open Science and the FAIR principles. It appears that the most relevant part of the UNESCO recommendation is Section 14a which gives the following motivation for open science: “Increased openness leads to increased transparency and trust […].”
I find it appropriate that a general declaration like this one does not prescribe a precise implementation. From a practical viewpoint, however, it remains somewhat unclear what one should write, due to our community’s inexperience in this regard. To give a picture of what current practice actually looks like, a list of examples of AI disclosure sections taken from recent arXiv preprints is added at the end of this blog post.
Additionally, journal policies indeed provide some guidance on the construction and motivation for such sections. For instance, the current journal policies of Elsevier state that “Authors should document their use of AI, including the name of the AI Tool used, the purpose of the use, and the extent of their oversight. Declaring the use of AI Tools supports transparency and trust between authors, readers, reviewers, editors and contributors and facilitates compliance with the terms of use of the relevant AI Tool. […]”.
One of the recommendations for mathematical institutions and funders is the following:
Align funding with values
Alignment with the values of this declaration should be taken into account in the evaluation and funding of projects which involve collaboration between academics and industrial partners.
– Leiden Declaration on Artificial Intelligence and Mathematics (2026)
I agree with this recommendation and would moreover remove the restrictive clause “which involve collaboration between academics and industrial partners”. Even the internal incentives in academia were already not fully aligned with our values. The ills of publish or perish and the theorem economy far predate this development. Similarly, the peer review process at journals was overloaded before AI entered the scene.
The present issues in the incentive landscape were also nicely summarized by John Baez: “I support this declaration. I have one small comment: the document notes that “Technologies which affect the way in which mathematics is practiced may disturb the current system of incentives.” The current system of incentives seriously is flawed in many ways, and I don’t think maintaining the status quo should be our goal. However, we should work to improve it, not let it be corrupted by outside forces, as has already been done for decades by university administrators, journal oligopolies, etc.”
It would have significant impact we could make some progress on changing the system of incentives in academia. This would have effects that exceed what can be accomplished by admonishing individuals to be responsible, and would be useful regardless of how AI develops into the future. Of course, changing the system is incredibly difficult. It is encouraging to see, however, that the number of signatories of the Leiden declaration is growing quickly and includes many signatories at advanced career stages like tenured professors.
Appendix: Examples of AI disclosure sections
Opening random papers on arXiv did not help me find many disclosure sections, since such sections are still quite rare. I hence used the “deep research” function on Google Gemini Pro 3.1 to find examples, which is where most of those below come from.
The following list is filtered for only those examples where the section primarily serves as disclosure and acknowledgement, as may appear in a research paper whose main topic is not the use of AI. There are also examples with much more extensive discussion on the uses and danger of AI, which is also interesting but something different. (e.g., arXiv:2510.23513, arXiv:2602.18918, arXiv:2604.07455, arXiv:2601.07222)
- (arXiv:2605.06697; math.NT) In the making of this paper, ChatGPT was used as a sounding board, for proofreading and for supplying numerical examples, most notably the calculation of the recurrence relation (4) in Section 3.2 and Table 1 in Section 5.3. The paper itself was entirely human-generated.
- (arXiv:2605.12030; math.CO) The proof of Theorem 1.1 is based on a suggestion from ChatGPT Pro. The final arguments in this paper are human-generated.
- (arXiv:2511.12799; quant-ph) AI-assisted tools (Claude, Anthropic) were used for code development, debugging, and manuscript preparation. All scientific concepts, experimental design, analysis, and interpretation were performed by the author. The author takes full intellectual responsibility for all content.
- (arXiv:2605.10478; math.AP) The authors used ChatGPT during the development of this work, including for suggesting proof strategies and assisting with calculations. All mathematical statements, proofs, and verifications were subsequently completed and rigorously checked by the authors, who take full responsibility for the results. We appreciate ChatGPT for its support.
- (arXiv:2603.21453; math.CA) Many of the references here were located by ChatGPT DeepResearch, Gemini DeepResearch, and Claude. The proof of Lemma 2.5 is based on suggestions from ChatGPT Pro, Gemini Pro, and Claude. GPT was used separately by Nat Sothanaphan, Aron Bhalla, and myself to identify corrections in an earlier version of the manuscript. As already mentioned, AlphaEvolve and ChatGPT Pro were used to discover the proof of Theorem 1.13 via Lemma 1.14. Github Copilot was used to perform autocompletion of text, Claude Code was used to perform routine typesetting, and Gemini was used to generate most of the plots in this paper; but besides those tool usages, the text in this paper was human-written.
- (arXiv:2605.00301, math.NT) ChatGPT was used to generate code for several of the images in this paper, to search for relevant literature (for instance, in locating references for the proof of Lemma 3.2), to proofread the paper and to offer additional suggested results and remarks11, and to perform numerics to guide the proof of Lemma 3.4.
The initial proof of Theorem 1.1 was generated by an autonomous run12 of GPT-5.4 Pro; a similar run also established Theorem 1.6. GPT-5.4 Pro was also used to assist with the initial proof of Theorem 1.2, with the main human contributions being the downward divisor chain and suggesting Lemmas 3.2 and 3.3(ii) to establish the sub-invariance property. In addition, an early version of GPT-5.5 Pro was used to assist with the initial proof of Theorem 1.3. Finally, GPT-5.4 Pro helped prove Theorem 1.4. Nevertheless, the final proofs in this paper have been generated and reviewed by the human authors, using the AI-generated proofs as starting points when appropriate.
The Lean formalization in [2] was generated using OpenAI’s Codex. The Lean formalization in [33] was generated using Math Inc.’s Gauss. - (arXiv:2606.02608, cs.LG) During the preparation of this work the author(s) used generative AI tools to accelerate drafting and revision. All mathematical claims, proofs, numerical interpretations, and bibliographic information were subsequently reviewed and edited by the author(s), who take full responsibility for the content of the manuscript.
- (arXiv:2509.26637, math.PR) During the preparation of this work, OpenAI’s ChatGPT was used to study unfamiliar proof techniques, suggest relevant references, and provide feedback on structure and clarity after the author had constructed the full arguments. It was also used during final manuscript preparation to assist with text refinement, grammar checking, and rewording for clarity and concision. All conceptual content, mathematical arguments, and original writing were developed by the author. The AI was not used to generate new ideas, mathematical structures, figures, or to interpret data. The author takes full responsibility for the correctness and final content of the manuscript.