We are running out of internet. For years, the recipe for better AI was simple: more compute, more parameters, and more scraped text. But if you want to build a model that understands the messy, high-stakes world of biology, Wikipedia and Reddit won't cut it. To actually engineer new proteins or predict how a drug interacts with a human cell, you need the kind of raw, proprietary data that usually stays locked behind the badge-access doors of corporate labs.
The Graveyard Strategy
OpenAI recently signaled a shift in its data acquisition strategy that every founder should pay attention to. Instead of just scraping public web data, they are looking at the intellectual wreckage of the biotech industry. The idea, originally championed by policy analyst Ruxandra Teslo, is both brilliant and a bit grim: bidding on the data silos of bankrupt biotech companies. When a startup goes bust, their lab notebooks, clinical trial failures, and proprietary chemical mappings usually gather dust or get sold for pennies on the dollar to liquidators who don't know what to do with them.
For a company like OpenAI, these bankruptcy proceedings are a goldmine. They represent years of human labor, millions in VC funding, and, most importantly, the negative results that never make it into academic journals. In science, knowing what didn't work is often just as valuable as knowing what did. By ingestion this data, AI models can learn the boundaries of biological possibility without having to repeat the expensive mistakes of the past.
Why General Models Fail at Science
The problem for builders in the AI space right now is the plateau of general intelligence. A model that can write a decent poem or summarize a meeting doesn't necessarily understand the nuances of protein folding or toxicity thresholds. Biology isn't a language in the way English is; it is a series of physical constraints and chemical reactions. General models are great at hallucinating plausible-sounding nonsense, but in medicine, a hallucination is a safety hazard.
OpenAI knows that to dominate the next frontier, they need specialized vertical data. They are paying to create or acquire high-fidelity biological datasets because the public web is essentially exhausted of high-quality scientific signal. This move validates a trend I've been watching: the shift from "quantity of tokens" to "quality of truth."
The Founder's Perspective
If you're building in the AI space, you need to ask yourself where your data edge comes from. If you are relying on the same Common Crawl datasets as everyone else, you don't have a moat; you have a commodity. OpenAI is showing us that the real value is moving toward closed, proprietary, and even "dead" data. They are essentially recycling the failures of the previous tech cycle to fuel the current one.
There is also a significant ethical and regulatory hurdle here. Biotech data isn't like a collection of cat photos. It involves sensitive patient information, manufacturing secrets, and potentially dual-use information that could be repurposed for bio-terrorism. While the bankruptcy strategy is clever, it bypasses the traditional gatekeepers of scientific publishing. We are entering an era where a private company could possess a more comprehensive understanding of human biology than any public university or government agency.
The Cost of Creation
Beyond buying up old data, OpenAI is also reportedly paying to generate new data from scratch. This is a massive shift in how these companies operate. They are no longer just passive observers of the web; they are active participants in the laboratory. By funding wet labs to run experiments specifically designed to fill the gaps in an AI's knowledge, they are closing the loop between digital prediction and physical reality.
For builders, this indicates a massive opportunity in "Data Synthesis Services." If you can create a pipeline that generates high-quality, verified biological or physical data that AI models are missing, you are holding the keys to the kingdom. The bottleneck is no longer the code; it is the ground truth.
Risk and Skepticism
I have to be honest: there is something slightly unsettling about the largest AI firm in the world picking through the bones of failed startups to train its next god-model. It raises questions about who owns the collective progress of science. If a startup took $50 million in investor money, failed to bring a drug to market, and then sold its data to OpenAI, does that data eventually become a public good, or is it locked forever in a proprietary black box?
Furthermore, we have to consider the "garbage in, garbage out" rule. Just because a company was well-funded doesn't mean its data is clean or useful. The technical debt in biotech is legendary. Integrating disparate datasets from dozens of different failed firms, each using different protocols and equipment, is a nightmare of a data engineering task. OpenAI is betting that their models can sift through the noise, but that is a very expensive bet to make.
The Takeaway
The low-hanging fruit of the internet is gone. The next phase of AI growth will be powered by the data we previously thought was useless or inaccessible. For builders, the lesson is clear: focus on proprietary data pipelines. Whether you are looking at bankruptcy courts, specialized sensors, or automated labs, the winner of the AI race won't just have the best algorithms—they'll have the best library of human failure and success.
The future of AI isn't just about reading what we've written; it's about learning from what we've done, especially the parts we'd rather forget.
If OpenAI succeeds in turning the biotech graveyard into a roadmap for new medicine, it will change the economics of failure in Silicon Valley. A failed startup might still be a successful data point. But until we see actual results—new drugs, new therapies, or new materials—it remains a high-priced experiment in corporate archeology.
Read the original at MIT Technology Review →