I have been watching the AI arms race from a founder's perspective for a long time, and the conversation usually settles into two camps: the closed-door giants like OpenAI and Google, and the open-weight advocates who want to give the power back to the individual developer. But lately, a new tension has emerged, and it is one that Garry Tan, the head of Y Combinator, is now shouting about from the rooftops.
The Distillation Gap
The problem is simple but dangerous for American builders. Right now, Chinese labs like Alibaba and DeepSeek are aggressively releasing open-weight models that are punching way above their weight class. They are doing this through a process called distillation. Essentially, they take a massive, expensive frontier model, use it to generate high-quality synthetic data, and then train a much smaller, leaner model on that data. The result is a compact model that performs nearly as well as the giants but runs on a fraction of the hardware.
Garry Tan’s argument, which he has been pushing lately, is that U.S. open-source labs need to start doing the exact same thing with American frontier models. If we don't, we risk a future where the only high-performance, accessible AI tools for startups are coming from overseas. For a founder, this is not just about geopolitics; it is about the sovereignty of your tech stack.
Why Builders Should Care
For most of us building products today, the dream isn't to pay a massive monthly toll to a closed API forever. We want to own our intelligence. We want to run models locally, fine-tune them on proprietary data without leaking that data to a third party, and scale without hitting rate limits that kill user experience. Open-weight models are the path to that freedom.
However, the current crop of U.S. open-weights often feels like a step behind. Meta’s Llama series is fantastic, but Tan is pointing toward a broader ecosystem where smaller labs can take the output of GPT-5 or Claude 3.5 and "distill" that intelligence into specialized, open tools. If American labs are restricted from doing this while international competitors do it freely, the competitive advantage for U.S. startups starts to evaporate.
The Synthetic Data Loop
Distillation relies heavily on synthetic data. This is where a large model acts as a teacher for a smaller student model. It creates a feedback loop that speeds up development significantly. Instead of scraping the entire internet and dealing with the messy, low-quality data found in the wild, distillation allows you to train on the "best hits" of a frontier model's reasoning capabilities.
Tan is effectively calling for a shift in how we view intellectual property and competition. If the goal is to keep the U.S. at the forefront of AI, we can't just have one or two winners behind a velvet rope. We need a robust middle class of open models that are distilled from the best frontier tech available.
The Risks of Falling Behind
If you look at what is happening in the open-source community right now, the most exciting benchmarks are often coming from models like Qwen or Yi. These are impressive feats of engineering, but they represent a strategic risk for American builders who might face future regulatory hurdles or supply chain issues if they become overly reliant on foreign infrastructure.
By encouraging U.S. labs to distill frontier models, Tan is advocating for a localized supply chain of intelligence. He wants the "distillery" to happen here, using models developed under U.S. safety standards and values, rather than ceding the open-source territory to labs that might have very different agendas.
Practical Reality for Founders
What does this mean for you? Right now, it means you should be looking at how you can use distillation in your own specific niche. You don't need to be a billion-dollar lab to benefit from this logic. You can use a frontier model to label a specific dataset for your industry—legal, medical, or creative—and then train a smaller, open-weight model to handle that task with high efficiency.
Tan’s vision is essentially a call to democratize the "brain" of the frontier models. If we allow distillation to become a standard practice for U.S. open-source labs, the cost of building an AI-first company drops significantly. You no longer need a massive compute budget to provide high-level reasoning to your users.
A Necessary Skepticism
We should be honest about the hurdles. There is a lot of legal gray area around using a competitor's model output to train your own. Most Terms of Service for the big labs explicitly forbid this. Tan is pushing for a world where this is not just tolerated, but encouraged for the sake of national competitiveness.
Whether the big labs like OpenAI or Anthropic will ever truly embrace this is unlikely. They have a massive incentive to keep their moats high. But the pressure from Y Combinator and the broader startup ecosystem is a signal that the "closed garden" approach is starting to frustrate the very people who are supposed to be building on top of these platforms.
The Takeaway
The future of AI isn't just about who has the biggest cluster of H100s. It is about who can most efficiently distribute intelligence. If Garry Tan gets his way, we will see a surge of American open-weight models that are leaner, faster, and just as smart as the giants. For builders, that means more choice, lower costs, and less reliance on a few gatekeepers in Silicon Valley or elsewhere. We need to stop worrying about just the frontier and start focusing on the distillation of that power into the hands of the many.
The goal is a robust set of open-weight options that ensure the U.S. remains the home of AI innovation, not just for the giants, but for the founders in the trenches.
Read the original at TechCrunch Venture →