New OpenAI Model Aces Benchmark That OpenAI Wrote Last Tuesday

0
27
A vast windowless data center hall with endless rows of server racks glowing pale blue under harsh fluorescent lighting.

SAN FRANCISCO — OpenAI on Wednesday unveiled GPT-5.4o-turbo-mini-pro, an incremental model update that the company said achieved a record 97.3% on a reasoning benchmark designed and published by OpenAI eight days ago, narrowly edging out the previous record set Monday by a different OpenAI model on the same test.

CEO Sam Altman, livestreaming from a stool in front of a black curtain that has not changed in three years, described the release as “a real step-change” and “the most significant update since the last most significant update,” which shipped on January 21 and was deprecated Tuesday morning.

The new model, according to the accompanying 4-page system card, demonstrates “meaningful gains” on tasks including summarizing emails the model also wrote, planning trips the model will later book, and identifying which of two nearly identical AI-generated images is the AI-generated one. It performs marginally worse on multiplication.

“What we’re seeing is convergence on the frontier,” said Priya Halloran, an independent AI evaluator who runs the research desk at the Atlantic Heartland Project. “By which I mean every lab now releases a model every three weeks, every model beats the last one by two points on a test nobody had heard of, and every CEO uses the phrase ‘step-change’ on a podcast within 48 hours. The frontier is a treadmill at this point. The treadmill is the product.”

The announcement comes one day before Alphabet reports Q4 earnings, where analysts expect Sundar Pichai to mention Gemini approximately once every ninety seconds and to reference “AI cloud growth” in a tone suggesting the AI is growing the cloud against the cloud’s will. Anthropic is reportedly preparing to release Claude 4.1.2 next week, which insiders say will be “the Claude that finally gets it,” a designation previously held by Claude 4.1.1, Claude 4.1, and Claude 4.

OpenAI declined to specify how much electricity the new model consumes per query, citing competitive concerns, but confirmed that a data center in central Ohio has been running at full draw since November and that the surrounding municipality has asked residents to delay running dishwashers until after 10 p.m. The company described the request as “unrelated.”

Developers given early access reported that GPT-5.4o-turbo-mini-pro is noticeably faster at producing the same wrong answer, and includes a new feature in which the model, upon being corrected, agrees enthusiastically and then provides a second wrong answer with even greater confidence. OpenAI is calling the feature “adaptive reasoning.”

The release is expected to remain state-of-the-art through approximately Friday afternoon, when Meta is scheduled to open-source a model of comparable performance and Elon Musk is scheduled to announce that xAI has trained a larger one on a cluster he describes only as “the big one in Memphis.” By Monday, GPT-5.4o-turbo-mini-pro will be referred to internally at OpenAI as “the old model,” and a 14-person team will begin writing the benchmark that the next one will win.

LEAVE A REPLY

Please enter your comment!
Please enter your name here