Loading prices…
STKR NewsSTKR News0 of 3 free this month
AI

Don’t be fooled by this summer of AI hype

Silicon Valley spent the summer shouting about AI security breakthroughs and model benchmarks, but behind the hype, builders are facing the same old reality: tools remain unreliable.

Originally on MIT Technology Review →
AB

Adrian Boysel

Contributor

Sep 22, 2026

5 min read

Photo illustration / STKR News

I spent the last few months watching the AI industry turn into a circus. If you have been following the news cycles from April through September, you have seen a barrage of claims that sound more like science fiction than software engineering. We are being told that the era of human error is over, that models are now superior to senior security researchers, and that we are one update away from perfect code. As someone who builds in this space, I can tell you: the reality on the ground looks nothing like the marketing decks.

The Vulnerability Mirage

It started with Anthropic claiming their Claude Mythos model could outperform the majority of cybersecurity experts at spotting software vulnerabilities. That is a massive claim. If true, it would mean we could fire half our security teams and let the API handle the audits. But builders know better. Finding a bug is easy; understanding the architectural implications of that bug and fixing it without breaking twelve other dependencies is where the real work happens. When a model claims to be 'better' than a human expert, it usually means it is better at a specific, narrow benchmark test that was likely included in its training data.

For founders, this creates a dangerous temptation to cut corners. If you trust the hype, you might start relying on these models for pre-production security checks. The reality is that these models often hallucinate vulnerabilities where none exist, or worse, miss subtle logic flaws that a human with context would catch in seconds. We are seeing a shift where the tools are getting louder, but not necessarily more accurate.

The Transparency Theater

Then we saw the string of disclosure incidents. We had the OpenAI-Hugging Face situation, followed by a wave of 'voluntary' disclosures from Anthropic and Meta. Anthropic wore their disclosure like a badge of honor, while Meta seemed to be dragged into the conversation kicking and screaming. This transparency theater is designed to make us feel safe, but it actually highlights how fragile these systems are. If these models were as secure and autonomous as the marketing suggests, we wouldn't be seeing constant reports of model weights leaking or prompt injections bypassing safety filters.

As a builder, this matters because it impacts your tech stack's integrity. When you integrate these models via API, you aren't just buying intelligence; you are inheriting the security debt of companies that are moving too fast to document their own failures properly. The 'summer of hype' was largely an attempt to distract us from the fact that the underlying infrastructure is still held together by digital duct tape.

Benchmarks vs. Building

Every week there is a new leaderboard. One model is 2% better at Python; another is 5% better at creative writing. But for those of us actually trying to ship products, these numbers are increasingly meaningless. We are seeing a plateau in actual utility despite the exponential increase in hype. The gap between what a model can do in a sterile testing environment and what it can do in a production environment with messy, real-world data is widening.

I have talked to dozens of founders who spent their summer trying to implement these 'breakthrough' features only to realize the latency, cost, and reliability issues make them non-viable for anything other than a demo. The industry is currently optimized for impressive demos, not for sustainable building. We are being sold a dream of effortless scaling while we are still stuck in the nightmare of prompt engineering and rate limits.

The Founder's Reality Check

If you are building in AI or crypto right now, you need to ignore the noise. The summer hype was a survival tactic for big labs looking to justify their next massive funding rounds. They need you to believe the models are getting smarter at an impossible rate so that the valuations keep climbing. But your customers don't care about Anthropic's latest benchmark score. They care if your app works, if their data is safe, and if your tool actually solves a problem.

Stop chasing the latest model version just because the press release says it's 'human-level.' Most of these claims are based on narrow datasets that don't reflect the complexity of a real business. Instead, focus on building defensive moats that don't rely entirely on a third-party API. If your entire value proposition is just a thin wrapper around a model that is allegedly 'smarter than a security expert,' you don't have a business; you have a temporary lease on someone else's marketing campaign.

Why the Sidetracking Matters

The danger of this hype cycle isn't just that it's annoying; it's that it sidetracks real innovation. When the industry focuses on who can shout the loudest about AGI, we stop solving the hard problems like data privacy, energy efficiency, and reliable output. We are seeing a massive brain drain as talented engineers move away from building useful tools and toward building 'hype-ware' designed to catch the eye of a VC.

We need to get back to the fundamentals of engineering. That means testing, validation, and honest assessments of what these tools can and cannot do. A model that can find a vulnerability in a 10-line snippet of code is a toy. A model that can understand a million-line codebase and provide actionable security patches is a tool. We are nowhere near the latter, despite what the summer headlines told you.

Takeaway for Builders

  • Verify, don't trust: Never take a model's security claims at face value. Run your own red-teaming and use human experts for anything mission-critical.
  • Ignore the leaderboards: Benchmark scores are the new vanity metrics. Focus on how the model performs on your specific edge cases.
  • Build for resilience: Assume the models will fail, leak, or hallucinate. Build your architecture to handle those failures gracefully.

The summer of hype is ending, and the cooling temperatures of reality are setting in. For the builders who stayed focused on utility rather than headlines, this is actually good news. The noise is clearing, and we can finally get back to work.


Read the original at MIT Technology Review →

The Brief

Stay Updated on Cutting-Edge Tech

A six-minute morning dispatch on the markets and the technology shaping them.

Free. No spam. Unsubscribe anytime.

Write for STKR

Become a Contributor

Earn $STKR for published stories on markets, protocols, and culture.

  • Earn $STKR for every published piece
  • Editorial support from the STKR desk
  • Byline visibility across the network
  • First look at the upcoming creator program
Apply to Write

Keep reading

All stories

Comments

24 reader responses