Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
The Navier–Stokes Millennium Prize Problem (simonwillison.net)
231 points by tosh 3 days ago | hide | past | favorite | 237 comments
 help



> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...

I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.

I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.


I have a similar story, but perhaps even stranger.

I work for a startup. We often bring a wooden arcade with us to conferences as a marketing gimmick.

The arcade runs a single side-scrolling video game. You're running from a monster and dodging obstacles. The goal is to survive as long as possible, and your result is measured in meters.

There are always a few competitive guys who spend the entire conference taking turns to play it. And every single time, the same thing happens.

Say the current high score is around 200m. Everybody fails somewhere around that number: 190m, 186m... Maybe someone manages 210m. And the high score moves up at a snail's pace.

Then, a new guy shows up and gets something like 500m on his third try. From their next turn on, everybody easily does 450 or more, even though they were struggling to get past 200 just one turn ago.

What makes it stranger is that the game is dead simple. It's not like the new guy discovered a move that unlocked this capability. And it wasn't a lack of motivation either - they'd all been playing for an hour already. They just started performing better after seeing it was possible. There has to be a name for this phenomenon.


Climbers know this phenomenon as the "send train". A group of climbers has been working on a problem for a while. When the first person sends it, it often happens that multiple people are successful in the immediate attempts that follow. Part of it is watching the successful beta, but it happens even when everyone knows the moves. Once you've seen it's possible, you stop climbing tentatively and commit to the hard move instead of hedging for a fall.

Reminds me of amateur table tennis. There is a strange dynamic where you down regulate your performance unconsciously when the opponent is playing worse and vice versa.

Could be described as some physical form of this effect: https://en.wikipedia.org/wiki/Asch_conformity_experiments

The term would be conformity / normative social influence.


Surely this has nothing to do with specifically table tennis, nor that it's amateur.

A better example would be mixed boys/girls sports classes in school, where the boys deliberately hold back as to not injure/scare the girls.

It's a pretty obvious and human thing not to go out and completely destroy a much weaker opponent. We're social animals after all.

There also may be an element of energy conservation, there's objectively no need to put in any more effort than necessary. Inefficient.


I am pretty sure if you play at pro level you train this out of yourself. But then, the marathon / running examples where certain milestones get broken by one person and then suddenly by a host of others, is a counterargument.

I think the difference is that, in the table tennis example, it happens unconsciously and you can't help it. But you're right, it likely happens in most sports.

The mixed boys/girls classes example is, as you said, deliberate. And I too remember holding back on purpose in such situations when I was a kid.

With the arcade, no one was deliberately holding back. I'm sure of it.


Is it that you’re holding back against weaker opponents, or that you’re more excited and engaged when you’re playing against someone that forces you closer to the edge of your ability?

I think it's largely due to psychology. If 210m is considered the best, then as they approach it they may start to tense up and choke. When the goal and possibility is known as 500m, then there's no point being concerned near 210m.

Achieving a high score requires focus and effort.

There's little reason to aim for a score a lot higher than the current best; and when getting close to the score you're aiming for, it's easy to get agitated and make a mistake.

And if you do beat the high score, you're likely to loosen your attention right after that, and it's even annoying to keep going for much more.


I like to leverage that when brainstorming solutions to hard problems. Instead of contemplating small percentage improvements, try to think about what's in the way of improvements that are orders of magnitude better (e.g., don't take time to run a big task from 100s to 90s, take it to milliseconds). Sometimes it unlocks big ideas.

I work in software deployment for large orgs. I use this technique to optimize. "How can I provision or deploy this with one package install, and launch instance." Typically take multi-page or multi-step deployments down to fully automated.

I wonder if it's the same with the sub 2 hour marathon which was broken this year (it certainly was with the 4 minute mile).

This is like the mile record. Everytime someone breaks the existing one a bunch of people surpass the old record shortly after.

Reference-dependent effort ("bunching" around a target) [1] is probably the closest thing. The existing high score was acting as everyone's reference point. Anchoring [2] is the more general cognitive version.

[1] https://www.nber.org/system/files/working_papers/w20343/w203... [2] https://en.wikipedia.org/wiki/Anchoring_effect


> I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.

If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.


Maybe that shows how scientific discoveries come to be. It's not a genius sitting alone in their chamber for a decade and then suddenly they emerge with this huge thing. That's Hollywood fiction. Scientific progress is the colaborative effort of countless researchers over long periods of time, communicating, exchanging ideas, many of them wrong, tweaking, trying, thinking, arguing. When a breakthrough happens then it's the tip of a mountain of work that came before it. If two individuals stand on that mountain and feel there is something somewhere then it's not too strange that they take the last step at roughly the same time because conditions were right. The preconditions were in place at that time, the results required for this were available and the focus was on this specific thing.

I'd say it illustrates well that this last piece, the person celebrated for the achievement, is disproportionally overvalued and the rest of the work they are standing on is disproportionally ignored.


This ties back into AI, too, right?

I always like to bring up how many decades of research, how many hundreds of years of entire PhD-theses, how many sleepless nights were used up to generate all the protein structure data that made up the corpus of Protein Data Bank - that was then hovered up by the AlphaFold team, and guess who got the Nobel Prize...


Individual crystal structures, while heroic, probably don't qualify for the Nobel Prize (the only counterexample that I can think of is the crystal structure of the ribosome, which finally demonstrated that the active site was RNA, not protein).

And don't forget that AlphaFold worked because of sequence data (similarly laboriously collected by a large number of people).

A big part of the problem is that we award prizes to the small number of people who do something first, or who were in charge- the LIGO NP should have been awarded to the entire consortium, not Kip Thorne and a couple other people under this logic.


Just like all non-AI related nobel prizes before them.

Every laureate stands on the shoulder of giants which is their field as a generations-spanning body of researchers. They did the last step and get recognized. The good ones acknowledge that in their acceptance speeches.


This seems to be the norm rather than the exception.

On a tangent, the genius of people like eg Einstein is not so much that he came up with all these things: other people were close, but that he was a singular individual that did all of these discoveries, instead of five different guys all making some breakthrough here or there.


If you believe the totality of the document, there was more shadiness in how OpenAI acted than just timing. Save other things they are accused of, the progression from rumors to replication would attract much less scrutiny. With those in mind timing begins to look suspicious at best.

Would anyone be surprised if major model companies had tagged the accounts of competitor employees for extra tracking? Given the concerns about distillation and bench marking it hardly seems irrational, but how it is used matters quite a lot.


You see this in many things. Look at climbing for instance, or running. Once the first v17 had been done, suddenly many people did them. Or running with a sub 2 hour marathon.

Not the hope, but telling them that an answer does exist. This ofcourse means now we are going to go into an even more darker cave next time we are looking for some gold and pathbreaking discoveries will become even more rare. Add to that, the fear of people not winning against ai and having fewer rewards, then fewer people even enter those fields or attempt problems over the next generation. AI erodes skills not at individual but at civilisational level.

It's something that happened before LLMs - multiple discovery. Calculus is a classic example.

Yes but in this case, the allegation is Leibniz literally looked into newtons notebooks

Yep - it's a different case. But the idea that without LLMs, it's unlikely the same idea would be discovered independently is not true - it does happen.

That's the other thing. Even if it somehow magically turns out that they didn't plagiarize Buckmaster and Alpöge, that an independent audit goes through all their processes and finds that they could not possibly have stolen anything, and ignoring their outrageous attempts to force a coauthor off the author list, and ignoring that their employees act like small children on social media, this is a company whose culture is “we heard of a breakthrough being possible, let's scoop it”.

Normally, when you tell a coworker that you're wrapping up a result, unless they're some kind of sociopath, their natural inclination would not be to try to steal it from you.


>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?

But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."


I’m not a mathematician so take this with a grain of salt. Apparently terry tao commented that the approach used for the Euler paper can “probably” be used for solving NS but it’s still technically challenging and can probably be done with an LLM with a lot of compute. To me the crux of the issue is whether the insight were stolen so that the problem becomes something that is in the domain of LLMs. This is much different than LLMs coming up with the insight. OpenAI wants everyone to think the LLM came up with the insight and solved the thing by itself even though they have perhaps an army of researchers.

Having the chat logs enter the training data and having them have a meaningful influence on the ultimate result the model produces are very different things.

The text for all the Goosebumps books are certainly in the training data and to some small amount influenced the solve. But their contribution was so vanishingly small it would seem absurd to say R L Stein should have recourse for contibuting to the solve.


But this is different, right?

The equivalent would be taking a (fully offline) LLM and asking it about the ending of one specific Goosebumps book, and it revealing the twist. And although that specific book was (probably) only once in the training data, a high parameter LLM can usually "remember" the twist.


The only way it would be able to tell you the ending is if it was somehow given more importance in pretraining, loaded into context, or represented in multiple sets of training samples. I have a blog that I make very LLM friendly and usually load posts up into context when I’m working on something relevant. I’ve also opted to improve models for everyone. Despite this, the model can’t recognize my site or any of my posts when I ask it to recall without internet usage (I also turn memory off btw).

That's not correct for a SOTA model with trillions of parameters. Those have immense amounts of knowledge trained into their parameters. Try it. It "remembers" the ending of random books, Goosebumps and otherwise.

That's the entire point of why knowledge cutoff is so important (if you use them offline).

I find it entirely believable that the Navier-Stokes conversation was auto-flagged as high value training data, and burned-in the models knowlegdge base.


It's almost as if it was actually found by manually written brute-force algorithm running on OpenAI's massive computer cluster.

I find the claims from OpenAI somehow more relatable and reasonable.

- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.

- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.

- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.

- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.


It’s fishy though that they heard one of seven problems was about to be solved and threw perhaps 15 million bucks at the right one.

[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]


Why in the world is that fishy?

Isn't that exactly what almost everyone would do given that they wanted to see how capable their model is and the tense competition they have with Anthropic right now? Stealing impressive headlines from your competitor is pure gold.


If they were willing to spend 105 million it, perhaps. But that would be surprising. It’s fishy because they spent something like 15 million on the right one.

Again, that is the point. There is rumours that this one thing would be solvable, so they focus on this, spending 15m on one specific thing, instead of 105m on many. How is that fishy?

I understood the rumors (as described by OpenAI) to be that “one of” the prizes was solvable. If they heard NS specifically was solvable and aren’t saying that, they’re intentionally obfuscating that.

Because that’s the one Anthropic was rumored to have solved.

"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."

That doesn't reject my claim. They just didn't name them in this post. It feels like, you're going through great lengths reading something into this.

No, sorry I was reading it correctly but they addressed the issue in the post itself. They make it clear they were aiming at all 7 problems looking for the "two" that were rumored to be solved. But they didn't go hog on NS until they made progress. See sibling comments. I misread it the first time to be a claim that they "heard one of 7 human intractable problems are solved and spent 15 million on the right one".

Didn't they say that they launched an effort to evaluate their model on all open Millennium Prize problems, and then narrowed down to Navier-Stokes as the most promising after seeing results on simplified versions of the problems?

Yes, you're right. It seems like they claim to have started broadly and narrowed it down based on some progress. That leads to different questions but does answer my initial "fishy" point.

LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.

If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?


This is also what I've been thinking. The result itself is amazing but it's not like this was completely unexpected. There has been a huge amount of progress on the problem in the last 10 years without which it seems unlikely today's full resolution would have been possible. It is not clear what strategy was taken but it sounds like it borrowed heavily from the two spanish mathematicians. Experts will scrutinize the proof and it will be interesting to see if anything truly original or unexpected was done, outside of known techniques, a move 37.