Loading prices…
STKR NewsSTKR News0 of 3 free this month
AI

OpenAI Says a Secret AI Model Cracked Hundreds of Open Math Problems in One Prompt—Mathematicians Want Receipts

OpenAI claims a secret model solved hundreds of math problems in one go. But without a public release or peer review, the math community is calling for proof over PR.

Originally on Decrypt →
AB

Adrian Boysel

Contributor

Oct 7, 2026

4 min read

Photo illustration / STKR News

OpenAI just dumped a mountain of mathematical manuscripts onto the internet, claiming their latest unreleased model solved hundreds of complex problems in a single prompt. If you have been following the AI space for more than ten minutes, you know the drill: big claims, impressive-looking PDF files, and a total lack of transparency regarding how the sausages were actually made.

As someone who talks to founders and builders every day, I find this strategy exhausting. We are moving out of the phase where "magic tricks" impress the market. We are entering the phase where we need to see the receipts. Right now, the mathematical community is asking for exactly that, and their skepticism is worth paying attention to.

The Single Prompt Myth

According to OpenAI, they fed a secret model a series of prompts that resulted in 722 manuscripts. They claim that for a significant portion of these, the AI cracked the solution on the very first try. For anyone who has spent time prompt engineering or building RAG systems, this sounds less like a technological breakthrough and more like a statistical anomaly or a very carefully curated dataset.

Math isn't like writing a blog post. In math, you are either right or you are wrong. There is no "hallucinating" a correct proof that holds up under scrutiny from a Fields Medalist. By claiming the model did this in one shot, OpenAI is trying to signal that they have achieved a level of logical reasoning that transcends the typical "next-token prediction" limitations of current LLMs.

But here is the catch: we do not know what the model is, we do not know how much reinforcement learning from human feedback was involved behind the scenes, and we certainly do not know if these problems were already present in the training data in some obscure corner of the web.

Why Builders Should Care

If you are building an AI startup, this move by OpenAI is a masterclass in narrative control. They are effectively staking a claim on "Reasoning" as their specific moat. By releasing 722 manuscripts without the model, they create a sense of FOMO among developers and researchers. They want you to believe that the gap between their proprietary tech and your open-source implementation is widening.

However, for a founder, the takeaway is different. If a model can solve a math problem in one prompt, that is a tool. But if a model requires ten prompts, a human-in-the-loop, and a verification layer to reach the same conclusion, that is a workflow. Workflows are where the real value is built. OpenAI is selling the dream of the autonomous genius, but the industry is still powered by the reality of iterative refinement.

The Transparency Gap

The academic and mathematical communities are not known for taking people at their word. They want the "receipts"—the logs, the parameters, and the ability to reproduce these results. OpenAI’s refusal to provide these details is a growing friction point. It creates a dynamic where the company functions more like a black-box research lab than a platform for builders.

When you build on top of a company that hides its best work behind a curtain of "safety" or "secrecy," you are building on shifting sand. We saw this with the transition from GPT-3.5 to GPT-4. The goalposts move, the pricing changes, and the "secret models" eventually arrive with caveats that were never mentioned in the initial hype cycle.

The Founder Perspective: Don't Chase the Ghost

My advice to builders is simple: ignore the secret models. If you can't access an API and run your own benchmarks, the technology doesn't exist for your business yet. OpenAI is playing a game of psychological warfare with competitors like Anthropic and Google. They need to stay in the headlines to justify their massive valuations and compute spend.

Real innovation in the math and logic space is happening in verifiable ways. Look at Lean and other formal verification languages. If AI can bridge the gap between natural language and formal logic, that is a game changer. But a PDF dump on a Friday afternoon isn't a breakthrough; it’s a press release.

The Skeptic's Corner

Why now? Why release these manuscripts today? Usually, these drops coincide with a competitor's release or a need to distract from internal turmoil. In the crypto world, we call this a "nothingburger" until the code is pushed to GitHub. In the AI world, we should treat unreleased models with the same level of healthy suspicion.

If this model is as good as they say, it shouldn't need a curated PR campaign to prove it. The results should be self-evident and reproducible by third parties. Until then, these 722 manuscripts are just expensive digital paperweights.

The Final Takeaway

OpenAI is claiming a massive win for AI reasoning, but the math community is rightfully holding the line on evidence. For founders, the lesson is to focus on what you can build today with the tools that are actually in your hands. A secret model that solves math problems doesn't help you ship code, acquire users, or solve real-world problems. Keep your eyes on the metrics that matter, and let the giants fight their PR wars in the background.

The value of a model isn't in what it can do in a vacuum; it's in what it can do for your users when the hype dies down.

Read the original at Decrypt →

The Brief

Stay Updated on Cutting-Edge Tech

A six-minute morning dispatch on the markets and the technology shaping them.

Free. No spam. Unsubscribe anytime.

Write for STKR

Become a Contributor

Earn $STKR for published stories on markets, protocols, and culture.

  • Earn $STKR for every published piece
  • Editorial support from the STKR desk
  • Byline visibility across the network
  • First look at the upcoming creator program
Apply to Write

Keep reading

All stories

Comments

24 reader responses