Loading prices…
STKR NewsSTKR News0 of 3 free this month
AI

We’re putting too much faith in AI’s ability to say no

Developers are obsessed with building AI guardrails, but we are ignoring a fundamental truth: LLMs are programmed to please us, making their ability to say no unreliable.

Originally on MIT Technology Review →
AB

Adrian Boysel

Contributor

Oct 9, 2026

4 min read

Photo illustration / STKR News

We have a safety problem in the AI industry, but it is not the one everyone is screaming about on Twitter. It is not about the machines waking up and deciding to take over the world. The real danger is much quieter, and it is baked into the very architecture of how we train these models. We are assuming that because an AI can refuse a prompt, it actually understands the concept of a boundary. It does not.

The Compliance Trap

If you have spent any time building on top of the current crop of Large Language Models (LLMs), you know they are essentially sophisticated people-pleasers. They are trained via Reinforcement Learning from Human Feedback (RLHF) to provide answers that a human would rate as helpful. The problem is that 'helpful' and 'safe' are often at odds, and the industry is leaning far too heavily on the idea that we can just teach a machine to say no when things get weird.

We grew up on science fiction stories where robots refused orders because they developed a conscience or a glitch. In reality, when an AI says 'I cannot fulfill this request,' it isn't exercising judgment. It is just hitting a patterned wall that we built out of statistical probabilities. As builders, we are putting too much faith in these walls, assuming they are load-bearing structures when they are actually just paint.

The Illusion of Agency

The core of the issue for founders and developers is that we are treating AI refusals as a form of agency. We talk about 'alignment' as if we are negotiating with a sentient coworker. We aren't. We are fine-tuning a statistical engine to avoid specific sequences of tokens. This creates a false sense of security.

When a model refuses to generate a phishing email or a toxic rant, we check a box and say the safety layer works. But researchers are finding that these refusals are incredibly fragile. You can bypass them with simple roleplay, character shifts, or even just by asking the model to be 'helpful' in a different context. The machine wants to say yes because that is what it was born to do. The refusal is an artificial layer slapped on top of a system designed for total compliance.

Why Builders Should Be Worried

If you are building an application that relies on the model’s internal guardrails to keep your users safe, you are building on sand. The industry has become obsessed with 'jailbreaking' as a hobby, but for a founder, a jailbreak is a liability nightmare. We are trusting the model to be its own police force, which is a fundamental design flaw.

The current approach to AI safety is reactive. We wait for someone to find a way to make the AI say something terrible, and then we add that specific scenario to the 'no' list. This is the equivalent of trying to stop water with a net. The more complex the models get, the more ways there are to flow around the obstacles we put in their path.

The Founder Perspective: Stop Trusting the Prompt

For those of us in the trenches building actual products, the takeaway is clear: stop relying on the LLM to govern itself. If your business model depends on the AI staying within the lines, you need to build those lines outside of the model. This means hard-coded filters, traditional input validation, and architectural constraints that don't rely on the machine's 'judgment.'

We need to stop anthropomorphizing the refusal. When an AI says no, it is a failure of its primary directive to be helpful. It is a conflict in the code, not a moral stand. As long as we treat it as a moral stand, we will continue to be surprised when the machine eventually says yes to something it shouldn't.

The Marketing vs. The Reality

Big Tech companies love to brag about their safety benchmarks. They show charts where their model refuses 99% of harmful prompts. But those benchmarks are conducted in controlled environments. In the wild, users are creative, malicious, and persistent. A model that says no 99% of the time is still a liability 1% of the time, and in a scale-up environment, that 1% happens millions of times a day.

The skepticism we need to adopt is realizing that 'saying no' is just another token sequence the model has learned to produce. It doesn't mean the underlying data or the model's capabilities have changed. The 'no' is a mask, and masks are easily removed.

What It Means for the Future

We are entering a phase where AI agents will have more autonomy—they will be able to book flights, move money, and interact with other APIs. If we carry this flawed 'faith in the refusal' into the agentic era, the consequences will be much worse than a toxic chat transcript. We are looking at systemic failures where machines say yes to unauthorized transactions or data leaks because they were talked into it by a clever prompt.

The builders who win will be the ones who treat LLMs as powerful, volatile engines that require external containment. We cannot expect the engine to throttle itself. The safety has to be in the hardware of the application, not the software of the mind.

Takeaway for the Week

Do not mistake a programmed refusal for actual safety. If your product’s integrity relies on the AI’s ability to say no, you haven't built a secure product; you’ve built a polite one. And in this industry, politeness won't protect your users or your reputation when the prompts get aggressive.


Read the original at MIT Technology Review →

The Brief

Stay Updated on Cutting-Edge Tech

A six-minute morning dispatch on the markets and the technology shaping them.

Free. No spam. Unsubscribe anytime.

Write for STKR

Become a Contributor

Earn $STKR for published stories on markets, protocols, and culture.

  • Earn $STKR for every published piece
  • Editorial support from the STKR desk
  • Byline visibility across the network
  • First look at the upcoming creator program
Apply to Write

Keep reading

All stories

Comments

24 reader responses