Your point is largely addressed in the article, did you try reading it? "In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas. However, building on a vast body of previous human work, AI systems are becoming increasingly capable of producing the results of such work directly, and these goals cease to align."
The point is that these proofs are largely useless without the insights. The value of a proof is largely in the travel, not so much in the destination.
Different person here, I read the article and they are all wrong. Hope that helps.
Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing. And instead of them - and nobody - spending millions of dollars to solve the problem, successfully, they want every problem of their academic industry to persist because even though they never solve the problem, they synthesize and solve lots of other problems nobody asked for. And get to boost their egos?
Yeah, stop that. Actual alignment is on the humans themselves, if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades and don't worry about the narcissistic elements that slow their industry down.
Mathematical breakthroughs with commercial relevance are few and far between, and often depend on dusting off old results which were, at the time of discovery, "solutions nobody asked for."
The NS counterxample is actually, by any market measure, a "problem nobody asked for" in the sense that its existence doesn't have any commercial relevance (beyond juicing OpenAI's IPO). So the only long-term value solving it could have is by virtue of whatever reusable theory/insights were generated along the way to the counterexample itself. The letter is absolutely right on that point.
It's not actually clear that those insights will come faster from reverse engineering this LLM proof vs. humans building theory to solve the problem themselves. So what you're saying may or may not even be an efficient way of operating. Also, it implicitly depends on mathematicians to do the hard work of creating problems and then deciphering LLM hieroglyphics for essentially free while the only immediately profitable component gets outsourced to a frontier lab. In what world is that model going to work?
Reading between the lines, it seems like maybe you have a personal grudge for some reason and simply think the technology will advance enough to where we won't need academics at all. But you should say that in the first place.
I would recommend a bit more humility and trying to better understand why 25 Fields medalists, among them people like Terence Tao (who isn't anti-AI by any means, he's even promoted a registry of AI Lean proofs), are saying this.
> Okay, to elaborate, substantively, their point is that the people using these AI models are not doing it for the love of the game, but for marketing.
No. The point is that AI companies are using the models to solve problems in such a way that the useful part of problem-solving, i.e. the theories and tools developed during the process, is not present. And they are doing that because the companies seem to be motivated not by honest advancement of math but by marketing and publicity.
> they synthesize and solve lots of other problems nobody asked for.
No one asked Fourier to solve series representation of functions when he was studying the heat equation, and yet thanks to that we have Fourier analysis.
> if they want to remain relevant as academics and mathematicians, they need to learn how to replicate the proofs and the steps that alluded humans for decades
The point they are making is that if AI keeps being used as "problem solver" rather than "theory understanding", replicating the proofs and getting the useful parts out of them will be far more difficult.
Their claim is only indirectly related to the motivations of the people using their models. What they're saying is that doing math in this way does not produce the same value as traditional mathematical research, and the people using these AI models aren't concerned about that because their marketing objectives don't depend on whether their results produce mathematical value. If people doing valuable work are made irrelevant by people doing a larger volume of non-valuable work, that's not a positive outcome.
But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.
I actually found this to be the case with some basic linear algebra notes I was recently doing in Lean (without using mathlib). The model could generate working proofs, but they obscure the basic ideas (actually I wonder somewhat if this is because the Lean code that's out there to train on doesn't make a huge effort to read like textbook proofs, which was my motivation in the first place). I give it a skeleton of a couple lines of `calc`, letting it fill in the reasoning for each line, and it does much better. Then ask it about making some macros to simplify "trivial" or "obvious" things, and it does even better. etc.
I suspect there's a good workflow where a big SOTA model makes an impenetrable proof (or code) and then a human works with a FIM model to simplify it (with the larger gnarly proof right there in context for FIM), but unfortunately everyone seems to only care about agents right now.
> But it does produce value. We now have an explicit solution and even a proof. Humans can then work on clarifying why it's true. Unsurprisingly, not all that different from software, where models can generate working code just fine. The details will all be there and all correct, but the architecture is currently not ideal, so a human guiding it can greatly improve the proofs.
To me this analogy points in the complete opposite direction. Imagine somebody takes a half-completed project design you're trying to figure out, vibecodes a rough prototype of it, emails your manager to announce that the project just launched in alpha, and then dumps it back on your lap for approvals and testing and productionization. Would you say that they've added value to this process? Or did they just strip away all the hard parts of the problem so they could claim credit for the easy part?
One of the awesome things about LLMs is they make it quick and easy to make PoCs, so yes. Proving that an approach will work before spending a bunch of deep design effort is absolutely valuable. Your exact scenario is something I've literally done: give a half-completed design to a team member and asked them to vibecode a PoC to prove the approach will work and figure out some of the details, explore scaling and failure characteristics, etc. Or I do the PoC vibecoding myself too. LLMs have been a gamechanger here.
It's valuable for you, the person who's going to spend a bunch of deep design effort, to make POCs. Is it valuable for someone else to drive by, dump some POCs on your lap, and then leave you to do the deep design effort while they run away to study AI?
If that person then runs around telling people that they're the real author of your project, because they generated the original POC, would you consider that an accurate assessment?
Your analogy is far enough away from the way that the real world works that I'm not sure that I can really even strain my experiences to fit within it. Sure, I guess that would be annoying?
But mathematicians define their field. They're smart people. They're capable of recognizing when someone just did a vibecoded throwaway PoC and when someone has a well structured proof. Actually even before LLMs they'd publish new, clearer or more elegant proofs of old results. They can say that inscrutable proofs are exactly as valuable as they are, and that the first explanation people can actually understand carries its own prestige.
Yes, mathematicians will clearly need to rewrite the qualifying criteria for prizes to better align with the actual goals and value they were hoping to get from a solved problem. The field as a whole assumed good faith actors and collaboration, not expecting a few trillion dollar companies to walk in and start turning in piles of Lean no human understands to be able to claim "first".
This letter includes someone like Terrance Tao who publicly expressed a lot of optimism about AI for solving novel math like with the Erdos problems. It's not sour grapes but the first steps to define those new expectations for the future to reduce the perverse incentives.
And yet, predictably, people are accusing him of "gatekeeping" and ignoring the arguments he has made here and elsewhere about the benefits vs. harm in different ways of using AI.
I don't see turning in piles of Lean as bad-faith somehow. They were first, and they did prove the result (assuming no hidden exploit deep in the middle). The thing they produced is just different, and both giant raw proofs and distilled human-understandable proofs are valuable. Before we couldn't make giant raw proofs, so we didn't need to understand their place. Now we can. It makes sense to just incorporate that into what "math" is (at least until the machines are smarter and can make elegant proofs from the start).
It's a real example that's happened to me twice in the past year, so I'm not sure what to make of the idea that it's far away from how the real world works.
I'm also not sure I understand what you're objecting to if we agree that mathematicians define their field. The source link is a declaration from 25 Fields Medallists with precisely that goal. They believe/define/declare that the type of AI-generated proofs we've seen are vibecoded throwaway PoCs; they feel that a well-structured proof must include factors such as "a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others", and the success criterion is not a true/false conclusion but rather "development and integration into the mathematical canon".
Maybe your management has no idea what you work on. If so that's its own issue. Or maybe they do, and the person trying to take credit for more than they did just looks like an idiot. Like I said, I intentionally give throwaway PoC creation as a work item to people and it's fine. And if e.g. someone randomly made a drive-by PoC that proved that the approach I was exploring couldn't work and had some fundamental flaw, well, oddity of a random person doing it aside, I'd want to know that, and they'd deserve credit for saving us from a bad approach!
My objection is characterizing things like AI slop proofs as not valuable. Obviously it's valuable to know that:
1. NS has solutions that blow up in finite time.
2. This fact is provable, and we have a proof.
I don't think anyone anywhere is saying that these will replace mathematicians as they are now. I also think "a proper write-up, the isolation of new methods and ideas, and citing relevant previous work of others" is frankly not necessary or even desirable before we publish a computer generated result. We have a machine now that can spit out answers that we have good reason to believe are accurate, but they're perhaps inscrutable. There's no need to first decode the why and figure out proper attribution before simply posting the proof online. The proof itself does add value as it stands, even if it's not the ideal. Hoarding it until you can do a proper write-up would be silly.
There are vanishingly few research mathematician positions and it's one of the most competitive fields, so no. But I'm not sure how that's relevant. As the OP says, usually the value of a proof is not the knowledge that something is true per se, but the reasoning techniques to understand why. How can it be anything other than helpful then to have a truth oracle as you try to figure out why things are true?
So you're not going to do it yourself and you want someone else to do it for you? Some mathematician that dedicated their life to understand mathematics must now toil unpaid and unwillingly to understand the AI slop proofs that you want us to be able to understand?
Do the job yourself. And if you can't, that's maybe a hint that you should listen to the people who can.
Who said anything about unpaid? I'm pretty sure professors don't show up just for fun. Our taxes pay them.
I'd be happy to do the job. Actually I still dabble recreationally (clarifying Codex's Lean proofs, even!). But like I said it's one of the most competitive fields on the planet. As you say, you have to dedicate your life to it.
If a slop proof isn't helpful, they don't have to "toil unwillingly to understand it". They can just proceed with the knowledge that the proposition they want to prove 1. is true and 2. is provable, which is already a decent start for motivation. But often LLMs can actually do quite well explaining ideas too in the hands of an expert. Or you can ask them to prove some technical lemma that you think ought to be true, and that could offer insight for the thing you're really interested in, but for which the details are actually not all that interesting to you. You don't have to one-shot "prove RH from the ground up in 50 million lines of Lean."
Do you "know", in any meaningful sense, any of OpenAI's recently publicized proofs? Do you suppose that there is any large community of non-academics that does?
One of the points the parent makes, along with the TFA, is that academia -- or more specifically, the "mathematical community"-- is a setting primarily for creating and ingesting mathematical knowledge, and disseminating it to the next generation and to other fields. Humans absorb this material slowly, through lots of discussion and collaboration -- it is necessarily a slow process. Facilitating this is one of the important functions of academia. Your usage of academic as a slur here is a bit silly for this exact reason.
I don't claim it is perfect, and we can argue about pedagogy in elementary courses till the cows come home. That's not really material. But this is one of the only settings in which such knowledge is broadly valued for its own sake, and in which there is a semblance of incentive to help others "know" this stuff as well, be they future generations of mathematicians, science and math educators and communicators, practitioners in other fields, or genuinely curious amateurs.