Skip to content
AI Side

88 Hours, 10,000 Agents: OpenAI Says It Solved Navier-Stokes: and It’s Already Turning Sour

Written by Gab

Contents

OpenAI claims to have produced a “solution” to the Navier-Stokes Millennium Prize Problem, thanks to a group of approximately 10,000 coordinated AI agents that worked for 88 hours. The company attributes the proof to an internal next-generation model described as “significantly more capable than GPT-6 Astra.”

The scope must be clearly established: OpenAI is announcing a solution and describing a research setup, but in this post it does not publish a complete proof, peer review, or independent validation.

Here is the post in which OpenAI makes this official claim.

The announcement is spectacular, but the debate is already no longer focused solely on the performance of a swarm of agents. In the replies, three questions are closely intertwined: the validity of a potential Navier-Stokes proof, the intellectual authorship of the ideas involved, and the conditions for accessing a research system that has remained closed.

What OpenAI claims, and what it does not claim

In its message, OpenAI says it is sharing a “solution” to the Navier-Stokes problem. The lab notes that this problem concerns whether a smooth description of three-dimensional fluid motion can “break down,” meaning lose its regularity. It also emphasizes that this question has remained open for approximately 90 years.

The central technical wording explicitly describes the claimed architecture:

"The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra." @OpenAI

Later in its post, OpenAI provides the figures underpinning the announcement:

"Our internal model group arrived at the Navier, Stokes solution in 88 hours, using around 10,000 coordinating AI agents." @OpenAI

The lab also states that it applied monitoring and isolation safeguards comparable to those used in its frontier model evaluations. This detail provides information about the operational framework claimed by OpenAI, but not about the mathematical content of the proof.

The key point is therefore this: OpenAI is claiming a solution produced by its internal system, not independent validation of that solution. The gap between these two stages is considerable. A proof addressing a Millennium Prize Problem must be read, tested, subjected to attempts to refute its lemmas, evaluated by specialists, and then potentially reviewed according to the criteria of the Clay Mathematics Institute.

Forced and unforced Navier-Stokes: the nuance that makes all the difference

A technical distinction changes the true scope of the announcement. The Millennium Prize Problem concerns the unforced Navier-Stokes equations, with no imposed external forcing term. The work surrounding this announcement, including that of Tristan Buckmaster and Levent Alpöge, concerns forced variants, in which a smooth external force is added to the equation.

According to reports published since the announcement, the result attributed to OpenAI’s internal model concerns a roughly hundred-page proof of blow-up for forced Navier-Stokes.

At this stage, no one has solved the unforced problem, the one carrying the Clay Mathematics Institute’s one-million-dollar prize.

This is the nuance that separates a genuine mathematical advance from a solved Millennium Prize Problem. It is almost systematically lost in media coverage.

The post does not name the model used. Nor does it say that it is “GPT-7,” contrary to @ziwenxu_’s interpretation. OpenAI refers only to a next-generation model that is significantly more capable than GPT-6 Astra. That is not the same as announcing a numbered version.

The message also does not specify the detailed role of the coordinated AI agents. Did they generate conjectures, explore the literature, propose lemmas, write the proof, perform computer formalization, verify logical dependencies, or combine all of these functions? The thread does not say.

The official announcement is not where this story began. A few hours earlier, Mark Kretschmann (@mark_k) posted a rumor that OpenAI might have solved Navier-Stokes and that an announcement was imminent. When asked for his source, his message pointed to a public statement by Tristan Buckmaster.

Here is the earlier post, brought into the OpenAI thread by @calebcterry.

This initial thread does not yet concern the 88 hours, the 10,000 agents, or OpenAI’s internal model. It is dominated by another uncertainty: who actually solved the problem, and in what order?

@TimTeaFan mentions a competing rumor involving Anthropic and its Fable system:

"Weren't there rumours that Anthropic solved it with Fable?" @TimTeaFan

@mark_k’s response is inconclusive:

"There is some dispute or drama going on right now." @mark_k

In response to a similar question about the identity of the person or people who solved it, he adds:

"Currently it's a bit unclear who exactly solved it. Maybe both jointly." @mark_k

The discrepancy between the two threads lies at the heart of the story. Before OpenAI’s announcement, the debate centered on a rumor, Anthropic, priority, and the identity of the people or labs involved. Following the announcement, OpenAI turns the rumor into a quantified, technical institutional claim, without addressing the questions of authorship and provenance that @mark_k’s thread had already raised.

The announcement therefore establishes one specific point: at this stage, OpenAI is the only organization to make a detailed official claim. On its own, it does not settle the question of scientific priority.

The Navier-Stokes equations are fundamental to fluid mechanics. They describe how a fluid, such as air, water, or a gas, evolves by relating several quantities:

  • the fluid’s velocity, represented by a field that varies in space and time;
  • pressure, which constrains and redistributes motion;
  • viscosity, meaning the fluid’s internal friction;
  • external forces, such as gravity or other applied forces;
  • the incompressibility constraint, in the classical case of an incompressible fluid.

In simplified form, the momentum equation is written as follows:

[ \partial_t u + (u \cdot \nabla)u = -\nabla p + \nu \Delta u + f ]

Here, (u) represents velocity, (p) pressure, (\nu) viscosity, and (f) external forces. For an incompressible fluid, the condition (\nabla \cdot u = 0) is generally added.

These few symbols encapsulate extremely complex dynamics. The nonlinear term ((u \cdot \nabla)u) expresses the fact that the fluid transports its own velocity. This is one of the main sources of difficulty: motion at one scale can feed motion at other scales, distorting and amplifying it.

The real question in three dimensions: global regularity or singularity?

The Millennium Problem does not simply ask whether a flow can be simulated. It poses a question of global regularity in three dimensions.

Given sufficiently smooth initial data, can it be proved that the solution remains smooth for all time? Or can it develop a singularity in finite time, a phenomenon often called blow-up?

A singularity does not necessarily mean that a real fluid physically “explodes.” In mathematical analysis, it means that certain quantities associated with the solution cease to be controllable within the required regularity framework. The derivatives of the velocity may, for example, become infinite or lose the properties needed to extend a smooth solution.

The difficulty can be summarized as follows:

  1. The initial data are regular.
  2. Viscosity tends to damp small structures and dissipate energy.
  3. But nonlinearity can concentrate activity at increasingly finer scales.
  4. One must prove that dissipation always prevails, or construct a case in which it is insufficient.

It is this global control over the balance between nonlinear concentration and viscous dissipation that has remained elusive for nearly nine decades.

Why have such compact equations resisted solution for so long?

The equations are short. Turbulence is not.

In turbulent flow, large-scale structures transfer energy to smaller structures. Vortices interact, deform, fragment, and recombine. This cascade of scales makes analytical estimates particularly challenging.

Mathematicians already have fundamental results. In particular, they know how to construct weak solutions in general frameworks and obtain stronger results in certain dimensions, for certain regimes, or over limited periods. But the open problem concerns the universal guarantee of smooth regularity in three dimensions, given smooth initial data.

That is why a potential solution would be a major achievement. It would not simply amount to having “better computed” a simulation, but to resolving a fundamental alternative concerning the possible behavior of these equations.

An announcement, even an official one, is not enough to formally solve a Millennium Prize Problem. The proof would need to be made available in a sufficiently complete form to allow thorough scrutiny by independent experts.

This notably entails:

  • publishing the full text of the proof and its assumptions;
  • making it possible to verify every lemma, every estimate, and every logical dependency;
  • critical review by specialists in partial differential equations and harmonic analysis;
  • resolving any objections raised during this evaluation;
  • allowing sufficient time to establish that the result is accepted by the mathematical community.

The Clay Mathematics Institute applies its own criteria for officially recognizing a solution to one of its Millennium Prize Problems. In particular, a proof must be published in a qualifying outlet, achieve general acceptance within the relevant community, and then withstand a period of scrutiny.

At this stage, OpenAI’s post provides neither a published proof, peer review, nor independent validation. It announces a claimed result whose mathematical recognition remains entirely to be established.

A Correct Proof Would Not Provide Perfect Weather Forecasts

One of the most intuitive, yet misleading, reactions is to directly link Navier-Stokes to exact weather prediction.

"Finally I’ll know for sure whether it’s going to rain tomorrow in London or not" @ChShersh

That conclusion does not follow from the announcement. Solving the theoretical regularity problem would not suddenly produce perfect weather forecasts.

Atmospheric forecasts depend notably on:

  • incomplete and noisy initial measurements;
  • coupled models that include humidity, clouds, radiation, and interactions between the ground and the atmosphere;
  • physical parameterizations;
  • necessarily discrete numerical computations;
  • the sensitivity to initial conditions inherent in chaotic systems.

A proof of regularity, if established, would clarify the mathematical structure of the equations. That would be a major achievement. But it would eliminate neither measurement uncertainty, modeling errors, nor operational computing limitations.

Likewise, it would not instantly transform aeronautical, naval, or energy engineering. These fields already use Navier-Stokes under specific assumptions, with simulations, turbulence models, and experimental calibration. A theoretical advance may enrich the available tools in the long term, but it does not automatically become an industrial product.

Scientific Credit, the Main Post’s Blind Spot

The most substantive response under the announcement does not only challenge the role of AI. It also concerns the visibility given to the human researchers who may have originated key elements.

@Quasilocal states the grievance directly:

"Ok but like want to mention Tristan Buckmaster and Levent Alpöge in the actual tweet where you claim credit, rather than in a separe quote-tweet not linked to from this one?" @Quasilocal

The criticism is specific. It is not merely about whether they were supposedly mentioned elsewhere. It targets the omission of Tristan Buckmaster and Levent Alpöge from the message that publicly claims the result.

OpenAI did not publicly respond to this challenge in the thread. However, @calebcterry extends this call for recognition by placing the result within a longer mathematical lineage:

"After doing more research on this, I think the human contribution is worth mentioning, too. My understanding is that this came out of years of work by Diego Córdoba and Luis Martínez-Zoroa developing the underlying approach, with Tristan Buckmaster and Levent Alpöge then using LLMs to push that research much further and eventually formalize and verify the result." @calebcterry

This statement should be read for what it is: an interpretation reported by @calebcterry in the thread, not a definitive attribution established by an academic publication reproduced here. Nevertheless, it raises the right question: what should receive credit when a system of agents produces a proof based on preexisting approaches, intuitions, formulations, and research?

@X_is_Arbitrary draws a sharper conclusion:

"This seems ballsy. If what the mathematicians claim is true, then this just bolsters claims that AI is incapable of new ideas. It needed two brilliant mathematicians to tell it how to solve the problem." @X_is_Arbitrary

This interpretation is not supported by the publicly available evidence. Based on the thread, it is impossible to determine precisely how the invention should be attributed among Diego Córdoba, Luis Martínez-Zoroa, Buckmaster, Alpöge, the LLM tools used by these researchers, and OpenAI’s agents. But the continuity between prior human research and the result claimed by the agents is a central issue, not merely a matter of communication.

OpenAI’s main post does not credit any named human researcher, while the replies make this omission the primary subject of public dispute.

Codex sessions: a serious, unsubstantiated allegation

The controversy takes a more serious turn with questions about possible Codex sessions associated with the researchers.

@Mmorgan_ML asks:

"Did you or did you not train on Buckmaster & Alpöge's Codex sessions?" @Mmorgan_ML

@melqartvantyrus essentially asks the same question. @ArtemR goes further by amplifying the allegation that research may have been stolen:

"But did OpenAI cheat by stealing the research?" @ArtemR

It is essential to distinguish facts from suspicions. These messages show that the allegation is circulating. They do not prove that Codex sessions, notes, private data, or work by Buckmaster and Alpöge were improperly used by OpenAI.

The exchanges reproduced here contain no response from OpenAI, no technical details about the relevant training data, and no corroborating evidence. Nor is the content of the external article cited by @ArtemR reproduced.

At this stage, the claim that Codex sessions were improperly used remains an unconfirmed allegation. Nevertheless, the question of the provenance of the data and research deserves a factual, public, and verifiable answer.

This issue is distinct from, though related to, the criticism concerning the system’s closed nature. Even in the absence of any established wrongdoing involving data, a lab with an internal model more powerful than publicly available tools creates a significant asymmetry in research.

What this announcement says about the state of the art

The most novel aspect of the OpenAI Navier-Stokes announcement is not merely the word “solution.” It is the claimed setup: approximately 10,000 coordinated AI agents, running for 88 hours, centered on an internal OpenAI model described as more capable than GPT-6 Astra.

This setup suggests a potential shift in how mathematical research is organized. Rather than a standalone conversational assistant, OpenAI describes a collective system capable of distributing tasks among subproblems, verification, proof attempts, searches for counterexamples, and synthesis.

But the industrialization of discovery raises a governance question. @CircumjovialLLC states it bluntly:

"Not OK to use private models trained on humanity's corpus of mathematics to solve open problems when those models are far more capable than what you released and are not accessible to all researchers. This is tech feudalism, leading to intellectual tyranny." @CircumjovialLLC

So far, no response from OpenAI addresses the release of the model, its weights, its compute cost, its energy footprint, or possible academic access. @Adamlags explicitly asks about the cost of the experiment, while @alono88 humorously asks about the token cost. These questions remain unanswered.

The criticism of restricted access therefore does not prove misconduct. It raises a structural problem: how can the community reproduce, audit, or challenge a discovery produced using a model and infrastructure to which it has no access?

Key takeaways and what remains to be established

OpenAI’s announcement is potentially historic. The lab claims to have found a solution to the Navier-Stokes problem in 88 hours using approximately 10,000 coordinated AI agents, with an internal model described as more powerful than GPT-6 Astra. This is a detailed official claim, whereas @mark_k’s earlier thread reported only a rumor and confusion over the possible authors.

But several critical questions remain open:

  1. Where is the full text of the Navier: Stokes proof?
  2. Which independent experts have reviewed it?
  3. Does it exactly meet the criteria of the Millennium Prize Problem?
  4. What precise role did the agents play: invention, exploration, drafting, formalization, or verification?
  5. How will OpenAI credit the contributions of Córdoba, Martínez-Zoroa, Buckmaster, Alpöge, and any other researchers?
  6. What data and materials were used to train or guide the internal model?
  7. Will questions about the Codex sessions receive a public, technical, and verifiable response?
  8. What access will the scientific community have to the results, methods, and computational conditions?

The real dividing line is not simply between humans and AI agents. It is between a spectacular claim and the scientific requirements of verification, reproducibility, and fair credit. OpenAI has provided figures and described a method. The thread highlights everything those figures do not yet allow us to determine.

Read in another language