OpenAI’s o3 model falls short of its own benchmark claims

Source Cryptopolitan

OpenAI’s newest LLM, o3, is facing scrutiny after independent tests found it solved a far fewer number of tough math problems than the company first claimed. 

When OpenAI unveiled o3 in December, executives said the model could answer “just over a fourth” of the problems in FrontierMath, a notoriously hard set of graduate‑level math puzzles.

The best competitor, they added, was stuck near 2%. “Today, all offerings out there have less than 2%,” Chief Research Officer Mark Chen said during o3 and o3 mini livestream. “We’re seeing, with o3 in aggressive test‑time compute settings, we’re able to get over 25%.”

TechCrunch reported that the result was obtained by OpenAI on a version of o3 that used more computing power than the model the company released last week.

On Friday, the research institute Epoch AI, which created FrontierMath, published its own score for the public o3.

Using an updated 290‑question edition of the benchmark, Epoch put the model at about 10%.

The result does match with a lower‑bound figure in OpenAI’s December technical paper, and Epoch cautioned that the discrepancy could be due to various reasons.

“The difference between our results and OpenAI’s might be due to OpenAI evaluating with a more powerful internal scaffold, using more test‑time computing, or because those results were run on a different subset of FrontierMath,” Epoch wrote.

FrontierMath is designed to measure progress toward advanced mathematical reasoning. The December 2024 public set contained 180 problems, while the February 2025 private update expanded the pool to 290.

Shifts in the question list and the amount of computing power allowed at test time can cause large swings in reported percentages.

OpenAI confirmed the public o3 model uses less compute than the demo version

Evidence that the commercial o3 is lacking also came from tests by the ARC Prize Foundation, which tried an earlier, larger build. The public release “is a different model… tuned for chat/product use,” ARC Price Foundation posted on X, adding that “all released o3 compute tiers are smaller than the version we benchmarked.”

OpenAI employee Wenda Zhou offered a similar explanation during a livestream last week. The production system, he said, was “more optimized for real‑world use cases” and speed. “We’ve done [optimizations] to make the model more cost efficient [and] more useful in general,” Zhou said, while acknowledging possible benchmark “disparities.”

Two smaller models from the company, o3‑mini‑high and the newly announced o4‑mini, already beat o3 on FrontierMath, and OpenAI says a better o3‑pro variant will arrive in the coming weeks.

Still, it shows how benchmark headlines can be misleading. In January, Epoch was criticized for delaying disclosure of OpenAI funding until after o3’s debut. More recently, Elon Musk’s startup xAI was accused of presenting charts that overstated the capabilities of its Grok 3 model.

Industry watchers say such benchmark controversies are becoming an occurrence in the AI industry as companies  race to capture headlines with new models.

Cryptopolitan Academy: Tired of market swings? Learn how DeFi can help you build steady passive income. Register Now

Disclaimer: For information purposes only. Past performance is not indicative of future results.
placeholder
Bitcoin CME gaps at $35,000, $27,000 and $21,000, which one gets filled first?Prioritize filling the $27,000 gap and even try higher.
Author  FXStreet
Aug 22, 2023
Prioritize filling the $27,000 gap and even try higher.
placeholder
Why a Quiet 2025 Signals a Massive 2026 Crypto Bull Run: Bitwise CIO ExplainsBitwise's Matt Hougan Predicts a Crypto Boom in 2026 Amid Current Market Struggles
Author  Mitrade
Nov 13, Thu
Bitwise's Matt Hougan Predicts a Crypto Boom in 2026 Amid Current Market Struggles
placeholder
Gold Price Forecast: XAU/USD recovers above $4,100, hawkish Fed might cap gainsGold price (XAU/USD) recovers some lost ground to near $4,105, snapping the two-day losing streak during the early European session on Friday. The precious metal edges higher on the softer US Dollar (USD).  Traders will take more cues from the Fedspeak later on Monday.
Author  FXStreet
22 hours ago
Gold price (XAU/USD) recovers some lost ground to near $4,105, snapping the two-day losing streak during the early European session on Friday. The precious metal edges higher on the softer US Dollar (USD).  Traders will take more cues from the Fedspeak later on Monday.
placeholder
Bitcoin slides deeper into red as bears lean on $96,600 wall and eye $90,000Bitcoin extends its decline after failing to reclaim $96,500, trading below $95,000, the 100-hour SMA and a bearish trend line near $96,600; unless bulls can force a decisive close back above $96,600–$97,200, the short-term path of least resistance stays lower, with $92,500, $90,000 and the main $88,500 support zone in focus.
Author  Mitrade
21 hours ago
Bitcoin extends its decline after failing to reclaim $96,500, trading below $95,000, the 100-hour SMA and a bearish trend line near $96,600; unless bulls can force a decisive close back above $96,600–$97,200, the short-term path of least resistance stays lower, with $92,500, $90,000 and the main $88,500 support zone in focus.
placeholder
Bitcoin briefly loses 2025 gains as crypto plunges over the weekend.Bitcoin experienced a sharp decline this weekend, briefly erasing its 2025 gains and dipping below its year-opening value of $93,507. The cryptocurrency fell to a low of $93,029 on Sunday, representing a 25% drop from its all-time high in October. Although it has rebounded slightly to around $94,209, the pressures on the market remain significant. The downturn occurred despite the reopening of the U.S. government on Thursday, which many had hoped would provide essential support for crypto markets. This year initially appeared promising for cryptocurrencies, particularly after the inauguration of President Donald Trump, who has established the most pro-crypto administration thus far. However, ongoing political tensions—including Trump's tariff strategies and the recent government shutdown, lasting a historic 43 days—have contributed to several rapid price pullbacks for Bitcoin throughout the year. Market dynamics are also being influenced by Bitcoin whales—investors holding large amounts of Bitcoin—who have been offloading portions of their assets, consequently stalling price rallies even as positive regulatory developments emerge. Despite these sell-offs, analysts from Glassnode argue that this behavior aligns with typical patterns seen among long-term investors during the concluding stages of bull markets, suggesting it is not indicative of a mass exodus. Notably, Bitcoin is not alone in its struggles, as Ethereum and Solana have also recorded declines of 7.95% and 28.3%, respectively, since the start of the year, while numerous altcoins have faced even steeper losses. Looking ahead, questions linger regarding the viability of the four-year cycle thesis, particularly given the increasing institutional support and regulatory frameworks now in place in the crypto landscape. Matt Hougan, chief investment officer at Bitwise, remains optimistic, suggesting a potential Bitcoin resurgence in 2026 driven by the “debasement trade” thesis and a broader trend toward increased adoption of stablecoins, tokenization, and decentralized finance. Hougan emphasized the soundness of the underlying fundamentals, pointing to a positive outlook for the sector in the longer term.
Author  Mitrade
21 hours ago
Bitcoin experienced a sharp decline this weekend, briefly erasing its 2025 gains and dipping below its year-opening value of $93,507. The cryptocurrency fell to a low of $93,029 on Sunday, representing a 25% drop from its all-time high in October. Although it has rebounded slightly to around $94,209, the pressures on the market remain significant. The downturn occurred despite the reopening of the U.S. government on Thursday, which many had hoped would provide essential support for crypto markets. This year initially appeared promising for cryptocurrencies, particularly after the inauguration of President Donald Trump, who has established the most pro-crypto administration thus far. However, ongoing political tensions—including Trump's tariff strategies and the recent government shutdown, lasting a historic 43 days—have contributed to several rapid price pullbacks for Bitcoin throughout the year. Market dynamics are also being influenced by Bitcoin whales—investors holding large amounts of Bitcoin—who have been offloading portions of their assets, consequently stalling price rallies even as positive regulatory developments emerge. Despite these sell-offs, analysts from Glassnode argue that this behavior aligns with typical patterns seen among long-term investors during the concluding stages of bull markets, suggesting it is not indicative of a mass exodus. Notably, Bitcoin is not alone in its struggles, as Ethereum and Solana have also recorded declines of 7.95% and 28.3%, respectively, since the start of the year, while numerous altcoins have faced even steeper losses. Looking ahead, questions linger regarding the viability of the four-year cycle thesis, particularly given the increasing institutional support and regulatory frameworks now in place in the crypto landscape. Matt Hougan, chief investment officer at Bitwise, remains optimistic, suggesting a potential Bitcoin resurgence in 2026 driven by the “debasement trade” thesis and a broader trend toward increased adoption of stablecoins, tokenization, and decentralized finance. Hougan emphasized the soundness of the underlying fundamentals, pointing to a positive outlook for the sector in the longer term.
goTop
quote