ResearchPapers, evidence & method
OpenAI says an unreleased model produced most of a set of math results from a single prompt to one agent
OpenAI told Scientific American that an unreleased model generated almost all of a set of math results from one prompt to one AI agent. Community commenters question what that framing means.
An OpenAI spokesperson told Scientific American that a model the company has not released produced almost every one of the results after a single prompt was given to a single AI agent. [1]
Commenters in the linked post argue that "one prompt to one agent" does not rule out that agent handing work to other instances. This is their guess, not a confirmed detail of how the system works. [1]
If the claim holds, it would suggest AI agents can work on mathematical research problems with little human direction.
Read the full assessment
How the system was set up, including whether it handed work to other agents, affects how much human or automated support the results really depended on.
"Solve math, Make no mistakes"² A single prompt handed to a single agent like it's worded doesn't mean the agent didn't spin up other agents. It was probably a prompt like "assess this list of open problems in mathematics, ones you deem plausibly solvable delegate to other instances to investigate and work on." Obviously a bit more to the prompt than that but no way this was one agent churning through every one of these problems First they came for the artists. Then they came for the mathematicians. Then they came for the agents, and I...