We have a tendency in the developer world to sprint toward standardization before we actually understand the blast radius of what we are building. The Model Context Protocol, or MCP, is the latest example of this rush. While it is being pitched as the universal bridge for agent-to-agent communication, a structural flaw in how it handles trust is turning that bridge into a highway for malicious exploits.
The Promise of the Universal Plug
For those who haven't been following the plumbing of the AI world, MCP was designed to solve a very real problem. If you build an AI agent, it usually lives in a silo. To make it useful, you have to write custom integrations for every data source—Google Drive, Slack, GitHub, or local databases. MCP was supposed to be the USB-C of AI: one protocol to connect any agent to any data source or another agent.
It sounds great on paper. Founders love it because it lowers the barrier to entry for building complex workflows. Instead of writing fifty API integrations, you write one MCP implementation and suddenly your agent can talk to everything. But the security community is starting to point out that when you build a bridge that connects everything, you also create a path for a fire to jump from one building to the next.
The Structural Trust Gap
The core issue identified by security researchers involves how prompts are passed and executed across these connections. In a standard software stack, we have clear boundaries. If I send a request to a database, the database doesn't suddenly start telling my server what to do. There is a hierarchy of control.
With MCP, these boundaries are blurred. Because the protocol is designed to be flexible and allow agents to share context, it creates a loophole where a malicious prompt can be injected from a low-trust environment into a high-trust environment. Think of it like this: your personal assistant agent, which has access to your calendar and email, asks a third-party research agent to look up a website. That website contains a hidden prompt that tells the research agent to tell the personal assistant to delete all your files. Because they share the same protocol and trust the "context" being passed back and forth, the attack works.
Why Builders Should Care
If you are building an AI startup right now, you are likely feeling the pressure to be "MCP compliant." Investors like it because it looks like scalability. Users like it because it looks like interoperability. But as a founder, you are the one who inherits the liability when a user's data gets wiped or leaked because your agent trusted a bad actor three steps down the chain.
This isn't just a theoretical bug that a patch can fix. It is a structural flaw in how we think about agentic workflows. We are treating agents like predictable software modules when they actually behave more like unpredictable human contractors. You wouldn't give a stranger the keys to your office just because they were introduced to you by a friend, yet that is essentially what the current implementation of MCP asks us to do.
The Reality of Indirect Injection
We’ve spent a lot of time talking about direct prompt injection—the stuff where you tell a chatbot to ignore its instructions. That’s easy to understand and relatively easy to guard against. What we are seeing with the vulnerabilities in Google's implementations and others using MCP is something much more dangerous: indirect injection.
In an indirect injection, the user isn't the one doing the attacking. The attack comes from the data the agent processes. If your agent is connected to the web or to shared documents via MCP, it is constantly ingesting potential instructions from outside your control. When those instructions are passed through a protocol that lacks strict sandboxing and verification, you've essentially given the internet a terminal window into your internal systems.
Rethinking the Agent Stack
So, what is the path forward for builders who actually want to ship secure products? It starts with extreme skepticism of any "universal" protocol that doesn't prioritize isolation. We need to move away from the idea that more connectivity is always better.
- Zero-Trust Architecture: Treat every piece of context coming through an MCP connection as hostile. Never assume that because a prompt came from an "authorized" agent, it is safe to execute.
- Human-in-the-loop: For any action that involves destructive capabilities or data exfiltration, there must be a hard stop that requires a human to click a button. You cannot automate trust.
- Sandboxed Environments: Agents should operate in ephemeral, locked-down containers. If an agent gets compromised via a malicious MCP prompt, the damage should be contained to that session, not the entire user account.
The Founder’s Perspective
I get the appeal of MCP. I really do. It feels like the early days of the web where we were just excited that things could talk to each other. But we are not building a hobbyist network anymore. We are building the infrastructure that people are going to use to run their businesses and manage their lives. We can't afford to be sloppy.
The rush to standardize is often driven by a desire to win a market, not to build a better product. If you are building in this space, your competitive advantage shouldn't just be that you connect to everything—it should be that you are the one people can actually trust. Right now, MCP is a red flag for anyone who takes security seriously.
The biggest risk in AI isn't that the models are too smart; it's that we are building connections that are too dumb.
We need to stop treating agents as trusted nodes. Until the protocol includes robust, cryptographic verification of intent and strict permissioning at every hop, it remains a massive liability. If you're integrating it today, you're not just building a feature; you're opening a back door and hoping nobody notices.
The Takeaway
Convenience is the enemy of security. MCP offers an incredible amount of convenience for developers, but the cost is a fundamental lack of control over how agents interact. For builders, the message is clear: don't let the hype of interoperability blind you to the basic principles of secure system design. Verify every input, sandbox every action, and never assume that a protocol will protect you from a malicious prompt.
Read the original at Ars Technica →