ÉTAT HER

OpenAI's 722 Math Papers: Single-Prompt Breakthroughs or Unverifiable Claims?

OpenAI published 722 manuscripts on GitHub citing progress on 372 difficult math problems, most generated by feeding one prompt to one AI agent. The underlying model remains undisclosed, and neither the prompts nor computation times were shared, raising doubts about verification.

OpenAI's 722 Math Papers: Single-Prompt Breakthroughs or Unverifiable Claims?

722 manuscripts, 372 hard math problems, with each result averaging roughly three hours of ChatGPT Pro usage—that's the scale of what OpenAI has just disclosed. But what's really intriguing is the method: according to what OpenAI told Scientific American, nearly every single paper was produced using one prompt handed to one AI agent.

All of the manuscripts have been uploaded to GitHub, verified by OpenAI's newly formed advisory group, the Advisory Group on Mathematics and Artificial Intelligence (AGMAI). Last month, OpenAI had already signaled its intentions, claiming its model "solved more than 100 long-standing open problems spanning most areas of mathematics." This release essentially lays out the concrete list. The company stated that the results were produced by a ChatGPT pilot model not yet released to the public.

Among the headline results OpenAI claims are a solution to the four-dimensional Kakeya conjecture, improvements to several key computer algorithms, and progress on the Riemann hypothesis—any one of which, if confirmed, would be a major event in the math world. In its announcement, the company wrote that it is starting with a GitHub repository release, complete with mechanisms for paper revisions and citations, "while still seeking other community-hosted options that meet the advisory group's guidelines," and promised that future releases will be more complete in terms of citation standards, mathematical exposition, and presentation.

But OpenAI only followed through on half of AGMAI's recommended transparency. The advisory group had originally recommended disclosing the compute time and prompts used for each individual problem, but OpenAI did not do this, instead only providing an overview of the overall reasoning process, compute estimates, and the number of problems attempted. In other words, the public can see the list of results, but cannot see the specific path that would let anyone verify how those results were actually produced.

This is exactly where the controversy lies. Prior to this, OpenAI's claims related to Navier-Stokes had already stirred up a wave of backlash in the math community, leaving many scholars wary of "one-shot AI problem solving" claims. MIT mathematician Andrew Sutherland told Scientific American: "Unless they release the model so results can be reproduced, any claim of a single agent solving a problem in one shot should be considered unverified."

The batch of manuscripts OpenAI has released still needs to be reviewed and verified by mathematicians one by one, and the real impact won't be clear until that review process concludes. As for whether the model itself will ever be made public, or when the specific prompts will be filled in, OpenAI has not given a timeline.

Related

What to know | $42B Net Loss and Another AI Extinction Warning—Anthropic Still Pushing for $2T IPO
Living

What to know | $42B Net Loss and Another AI Extinction Warning—Anthropic Still Pushing for $2T IPO

The key details: Anthropic's IPO draft rarely admits in its own prospectus that its AI could pose an "existential risk" to humanity, while also revealing a $42 billion loss last year—yet its valuation could still reach $2 trillion.

Ring's Fix for Smart Lock Skeptics: A Manual Crank When Power Runs Out
Living

Ring's Fix for Smart Lock Skeptics: A Manual Crank When Power Runs Out

According to Ring founder Jamie Siminoff, fear of dead batteries has long held smart locks back. The company's answer is a hand-crank generator built into its upcoming Smart Lock, letting users turn a handle for emergency entry. Launch is set for early 2027 at $249.

Your Google One Tier Now Decides Access: Google Launches Experimental Game-Building Platform Playground, Just Type to Create
Living

Your Google One Tier Now Decides Access: Google Launches Experimental Game-Building Platform Playground, Just Type to Create

Access to Google's new AI game creation tool, Playground, is tiered according to your Google One membership level, and for now limited to US users aged 18 and older. The platform lets anyone generate 2D or 3D games from a few text prompts alone, without coding or drawing skills.

HBO Max, Paramount+ and Discovery+ to Merge Under Skydance After $76B Deal Closes
Living

HBO Max, Paramount+ and Discovery+ to Merge Under Skydance After $76B Deal Closes

Skydance confirmed its HBO Max, Paramount+ and Discovery+ platforms will be gradually folded into a single streaming service. The move accompanies the completed $110 billion Warner Bros. Discovery acquisition, with the merged service's name and pricing still undecided.

It's Official: Paramount and Warner Bros. Discovery Complete $11B Merger Under Skydance Banner
Living

It's Official: Paramount and Warner Bros. Discovery Complete $11B Merger Under Skydance Banner

Nearly a year after talks began, regulators have signed off on the deal. The combined entity will operate as Skydance — now carrying $80 billion in debt as part of the arrangement.

25 Years Later, PS2's Original SPC970 Security Chip Finally Gives Up Its Secrets
Living

25 Years Later, PS2's Original SPC970 Security Chip Finally Gives Up Its Secrets

After a four-year effort, developer DiscoStarslayer and collaborator Libby exploited an EEPROM write flaw to extract firmware from the original PS2's MechaCon security chip, posting 22 image files to GitHub.