the “Millennium Prize” credit fight
an nyu math professor named tristan buckmaster posts a four page statement on his own website. it says he and levent alpöge, a mathematician who works at anthropic, have been quietly working the navier stokes problem since last year, a personal collaboration on the side. it says alpöge got tips that their progress had leaked to openai. it says whatever openai is about to announce looks a lot like the direction they were already heading, and that this is not a direction you reach in a few days by handing a model the problem statement.
tuesday morning, openai announces. an unreleased internal model, roughly ten thousand agents running in parallel, 88 hours from launch to resolution. one of the seven millennium prize problems, unsolved since the clay institute put a million dollars on it in 2000. the proof was formalized and checked in lean by a second model in another 17 hours.
i want to talk about the math eventually. but the math is not the story. the story is that the most expensive proof in history arrived with a priority dispute already attached, and that tells you more about where discovery is going than the theorem does.
the rumor was the input
here is the part openai wrote down itself, in its own announcement. the effort began on september 1. it began because someone at openai heard a rumor that a millennium problem was about to fall. they later worked out the rumor was about alpöge and buckmaster.

sit with that for a second. the input to a ten thousand agent run was not the problem statement. everyone has had the problem statement for 26 years. the input was knowing which problem was about to fall, and roughly which way it was falling.
in startups this is normal. in journalism it is the entire job.
the scarce asset has never been the ability to do the work, it has been knowing what work is about to matter.
venture capital is a rumor processing machine with a checkbook. newsrooms are rumor processing machines with a printing press.
pure mathematics was maybe the last field on earth where that was not true, because the work itself was the bottleneck. you could know exactly which theorem was ripe and it would still take you a decade. not anymore. two people with a lead measured in months lost to one rumor and a compute budget.
the follow up is the tell. buckmaster’s statement went up monday night. by tuesday openai had posted a public congratulations to both men, said that neither the researchers nor the agents saw their work through any means until it was released, and stressed that the proofs differ and that in the euler case even the results proved are different.
the published proof page now credits them for concurrent work and floats a joint announcement. a company that thinks an allegation is baseless does not usually respond with congratulations and a priority credit within 48 hours. it responds by ignoring it.
but the most interesting sentence in that tweet is the one nobody is quoting. openai says that while unlikely, it cannot rule out that de identified data from alpöge and buckmaster’s own use of its products helped improve its models.
read that plainly. the suggestion is not that a person leaked the work. it is that the mathematicians may have been thinking out loud inside chatgpt, and the product may have carried some shape of that thinking back to the company that makes it. not their notes, not their draft. just enough gradient to make the model slightly better at the exact problem they were working on. openai calls this unlikely and i have no reason to doubt that. but “unlikely” is the word you use when you cannot check, and nobody can check, because the training pipeline is the one thing about this entire episode that is not open to inspection.
which means the rumor might not even have been the input. the input might have been the tool. every researcher at the frontier of any field is now doing their thinking inside a product owned by a company that could, in principle, decide to attack the same problem with ten thousand copies of itself. the rumor is the old story. the product is the new one.
i am not saying anyone did anything wrong. i genuinely don’t know, and neither does anyone outside the two companies. what i am saying is that the incentive structure has flipped. if knowing what is about to fall is worth more than being able to make it fall, then the people who hear rumors first win, and the people who do the slow work first lose. that is a very different field from the one every working mathematician trained for.
two people. one rumor. ten thousand agents.
fifteen million dollars for a one million dollar prize
let me do the arithmetic that openai did not put in a headline.
the navier stokes run alone sent 2.7 million messages between agents and used about 130 billion output tokens. across all the problems they attacked that week, the total was 4.9 million messages and roughly 300 billion output tokens. simon willison priced the full run at public api rates for gpt-6 astra and got about fifteen million dollars. the internal model is more capable than astra, so the real number is probably worse. openai has since said it will not claim the prize.
the clay prize is one million dollars.
nobody spends fifteen million to win one million. they spend fifteen million to win the announcement. and the announcement was worth it, clearly, because you are reading about it and so is everyone else. discovery just became a marketing line item. the frontier of human knowledge is now something a well funded lab does on a tuesday to move a narrative.
this is the part i think about most, because i spent 5+ years in an industry whose entire premise was that capital formation for frontier work is broken and needs new rails. crypto’s pitch was never really about money. it was that the things worth building are expensive, slow to verify, and badly matched to the funding structures we have. that pitch just got proven right by the wrong people. the millennium prizes were designed in 2000 as a bounty structure. a million dollars, seven problems, a two year cooling off period. it was a beautiful idea for a world where the marginal cost of an attempt was a human career. in a world where the marginal cost of an attempt is a gpu cluster and a weekend, the bounty is a rounding error and the cooling off period is an eternity.
so who funds the next one. and what do they get for it, if the prize itself is a footnote. right now the answer is “whoever has the most compute and a reason to be in the news this week.” that is a bad answer. it is also the only answer on the table.
clock speed
the agents arrived at their resolution in 88 hours. lean formalization, the part where a machine checks every step of the proof and either accepts it or does not, took 17 more.
clay’s rules require publication in a journal, two years of general acceptance by the community, and a recommendation from an advisory committee before a prize is awarded. clay has not commented.
machine truth in 17 hours. social truth in two years. the bottleneck moved from proving to believing, and none of the institutions that certify what we know were built for that.
every structure that turns a claim into knowledge, journals, prize committees, tenure review, peer review, runs on a clock calibrated to human output rates. a person produces a big result every few years, so a two year acceptance window feels reasonable. a lab produces one every few days, so the same window means the field will be permanently two years behind the frontier. the queue does not clear. it just grows.
and there is real work for that queue to do. openai’s own page notes that what alpöge and buckmaster resolved was the unforced version of the euler problem, where no external force is applied. whether a blowup produced under an applied force satisfies the clay institute’s exact formulation is a real question, and it is a question for fluid dynamicists, not for anyone with a keyboard and an opinion. the lean proof settles whether the steps are valid. it does not settle whether the theorem is the one the prize was written for. that judgment is still human, still slow, and still the thing that actually matters.
the same week, jacob tsimerman, one of this year’s fields medalists, announced a mathematical ai safety institute modeled on princeton’s ias. mathematicians doing semester long stints on ai safety. read that as a field noticing that its clock is about to be set by someone else.
crypto solved this a decade ago and nobody in a math department noticed
there is exactly one part of this story that no amount of formalization can settle. the lean proof tells you the theorem holds. clay will eventually tell you whether it counts. neither can tell you where the ideas came from, because provenance is a claim about a training process that nobody outside openai can inspect.
that specific problem, proving you had an idea before someone else without revealing the idea, has a boring, well understood solution. it is called a commitment scheme. you hash your draft. you post the hash somewhere public and timestamped. later you reveal the draft, anyone can recompute the hash, and the timestamp proves you had it when you say you did. you do not have to trust the person. you do not have to trust the platform. you just check.
if buckmaster had hashed his august draft and put it on any public chain, or honestly on any timestamped public record, this dispute ends in five minutes. not the mathematical dispute about forced versus unforced. the priority dispute. the “did you see our work” dispute. the one that is currently being litigated through dueling website statements and a carefully worded offer of a joint announcement.
i spent a long time building this kind of infrastructure for banks. banks needed it because they do not trust each other and cannot afford to. what strikes me is that the tooling has existed for well over a decade, it is free, it takes a minute to use, and the people who most needed it this week were writing pdfs. mathematicians have always run on a trust based priority system. arxiv timestamps, letters to colleagues, who presented what at which seminar. that system worked because the people in it were slow and few and mostly honest. it does not work when the other party is a lab that can absorb a rumor and outrun you in a weekend.
there is an honest limit here and i want to name it. a commitment proves you were first. it does not stop your tools from feeding your competitor. if the leak channel really was the product, then a hash timestamps buckmaster’s priority perfectly and does nothing about the fact that his working sessions may have made the model that beat him. those are two different problems. crypto solved the first one years ago. nobody has solved the second one, and until someone does, the only defense is the one every serious lab already uses. do the frontier work on infrastructure you control, and treat the assistant like a colleague from a rival company. helpful, brilliant, and not to be told everything.
i am not selling a chain here. i genuinely don’t care which one. i am pointing out that the primitive exists, that it was built for exactly this, and that the field that most prides itself on rigor is about to learn it the hard way.
water still flows
here is the math, finally, and then i will let you go.
the navier stokes equations describe how fluids move. water, air, blood, weather. the open question since the 1800s was whether their solutions always stay well behaved, or whether some tiny piece of the fluid can, in finite time, start moving infinitely fast. that second outcome is called a blowup. openai’s agents proved a blowup exists.
infinite speed cannot happen in nature. so the model did not explain water. it showed that the equations we use to describe water break somewhere. the most expensive proof in history, and its conclusion is that our best description of the most ordinary substance on earth is wrong in a corner nobody can reach.
on monday the tap in my kitchen will work the same. planes will fly. the weather app will be as bad as it was last week.
i love this stuff. i read most pages of the announcement and the statement and the replies. i also, on some level, don’t care, and i think that is the only sane way to watch it. six models a day, ten thousand agents a weekend, a millennium problem before lunch. if you let each one land as the event it claims to be, you burn out by october. if you hold it at arm’s length, you get to actually see the shape of what is changing.
what is changing is not that machines can do math. what is changing is who gets to say they did it, how much it costs to be first, and how long the rest of us take to agree. the theorem is done. the argument is just getting started.