The Real Cost of Wrapping Models
In the early days of any gold rush, the people selling the picks and shovels are the ones getting rich. In the AI world, we assumed those people were the model providers. We thought the margins of a startup would be dictated by the price per token set by OpenAI or Anthropic. We were wrong. It turns out the biggest factor in your profitability isn't what the model charges you, but how you built the cage around it.
A recent study coming out of UC Berkeley sheds light on a phenomenon that founders need to pay attention to immediately. They looked at how different 'harnesses'—the software layers that wrap around an AI model to perform specific tasks—affect the cost and quality of the output. The results were jarring. If you are building an AI company today, you are likely either bleeding money through inefficient engineering or sitting on a goldmine of margin you haven't tapped yet.
The Seventy-One Percent Difference
The researchers looked at a specific model, GPT-5.6 Sol, and tested its performance across two different coding platforms: Pi and Claude Code. Both platforms use the exact same underlying model to solve the same problems. You would assume the cost to the user or the cost to the company would be relatively similar. Instead, the study found that the cost on Pi was 71% lower than on Claude Code.
Let that sink in for a second. We are talking about the same brain returning the same result. The only difference is the harness. When we talk about a 'harness,' we are talking about the prompt engineering, the retrieval-augmented generation (RAG) pipelines, the memory management, and the way the software interacts with the API. A 71% difference in cost for the same quality of output is the difference between a sustainable business and a failed experiment.
What is even more interesting is that across 42 different harness comparisons, the researchers found no statistically significant difference in quality. The cheaper versions weren't 'worse.' They were just smarter. They used fewer tokens, smarter caching, or better context management to get to the finish line.
The Tale of Two Startups
To understand what this means for founders, let's look at the math behind a hypothetical $250,000 contract. Imagine two startups bidding for the same enterprise deal. Both are using the same foundational models. Startup A has built a heavy, unoptimized harness. Startup B has spent the time to build a 'factory' for automating their performance improvements.
Startup A might end up with a 38% gross margin. After they pay for the compute and the tokens required to fulfill that contract, they are left with a decent chunk of change, but their customer acquisition cost (CAC) payback period is going to be long. They have to support a heavy infra load for every new user they onboard.
Startup B, using a leaner, more efficient harness, could see gross margins as high as 75%. Because their cost to serve is so much lower, they pay back their sales costs in half the time it takes Startup A. In a venture-backed world where 'efficiency' is the new 'growth at all costs,' Startup B is the one that gets funded. Startup A is the one that gets acqui-hired or shuts down.
Why We Overspend on Tokens
Most builders are lazy when it comes to token consumption because they are focused on shipping. We throw the kitchen sink at the prompt. We send the entire documentation folder into the context window because we can. But the Berkeley study proves that this 'brute force' approach to AI is a margin killer.
The reason one harness is cheaper than another usually comes down to three things: customer understanding, relevant evaluations, and automated hill climbing. If you don't know exactly what your customer needs, you over-prompt. If you don't have a way to measure quality (evals), you can't trim the fat without fearing the model will break. And if you aren't automating the process of testing new, cheaper ways to get the same answer, you are stuck with your first, most expensive draft.
Building a Performance Factory
If you want to survive the next two years of the AI cycle, you have to stop thinking like a prompt engineer and start thinking like a manufacturing engineer. You need to build a factory that automates the optimization of your harness. This means constantly running A/B tests on your prompts, your RAG retrieval chunks, and your model routing.
The harness sets the price of the answer, not the model. If you control the harness, you control your destiny.
We are entering an era where 'LLM Ops' isn't just a buzzword; it is a financial necessity. Founders who treat their AI costs as a fixed commodity are going to be disrupted by builders who treat their AI costs as a variable they can aggressively optimize through better software architecture.
What This Means for Builders
The takeaway here is clear: stop obsessing over which model is 2% better on a benchmark and start obsessing over how efficiently your software uses those models. The margin opportunity in AI isn't in building the best model; it's in building the best harness.
- Audit your token spend: Look at your most expensive queries and see if a smaller context window or a more refined prompt produces the same result.
- Invest in Evals: You cannot optimize what you cannot measure. Build a robust evaluation suite so you can aggressively cut costs without losing sleep over quality drops.
- Focus on the Harness: Treat your software wrapper as a core product feature, not just a middleman between the user and the API.
The startups that win won't be the ones with the most tokens; they will be the ones that need the fewest tokens to deliver the most value. That is how you build a real business in this space.
Read the original at Tomasz Tunguz →