I have seen enough AI hype cycles to be naturally suspicious when a new lab claims they have suddenly jumped over OpenAI and Anthropic. Usually, it is a cherry-picked benchmark or a very narrow use case that doesn't hold up in the wild. But when the team behind the claim is composed of DeepMind alumni, you have to at least stop and look at the architecture. That is where we are with Inherent and their new agent, Faraday.
The Replication Crisis Meets the Compute Boom
For those of you who haven't spent time in the trenches of academic research, replication is the ultimate vibe check. A researcher publishes a paper claiming a breakthrough; another researcher has to take that paper, follow the steps, and get the same result. It sounds simple, but in practice, it is a nightmare. Documentation is often incomplete, environments are hard to set up, and small variables lead to massive deviations.
Inherent claims that their AI agent, Faraday, just outperformed Claude 3.5 Sonnet and GPT-4o at this specific task. We aren't just talking about summarizing a PDF. We are talking about an agent that can ingest a paper, understand the methodology, write the code, and actually execute it to verify the findings. This is what we call an 'autonomous research teammate,' and if the data holds up, it marks a shift from LLMs as assistants to LLMs as contributors.
Why Generalist Models Struggle With Science
Most of the models we use every day are generalists. They are great at writing emails, coding basic CRUD apps, or hallucinating historical facts. But scientific research is a different beast. It requires a level of logical persistence that most current models lack. If a model hits a bug while trying to replicate a study, it often loops or gives up. It lacks the 'founder mindset' of troubleshooting until the job is done.
Inherent is approaching this differently. By focusing specifically on the scientific workflow, they are building for depth rather than breadth. For builders, this is a lesson in niching down. While everyone else is trying to build the next 'everything app' for AI, the real value is being captured by teams solving hard, high-friction problems like peer review and experimental verification.
The Architecture of an AI Teammate
What makes an 'agent' different from a 'chatbot' is the ability to operate in a loop without a human holding its hand every five seconds. Faraday is designed to function as a teammate. This means it has access to tools, it can run simulations, and it can self-correct when the initial output doesn't match the expected parameters of the research paper it is analyzing.
The Inherent team, leveraging their experience from the early days of DeepMind, seems to be focusing on the 'reasoning' layer rather than just the 'prediction' layer. In my experience building in this space, the biggest hurdle for AI agents isn't intelligence—it is reliability. If an agent can accurately replicate a complex study, it proves a level of reliability that could eventually translate to autonomous engineering and drug discovery.
What This Means for the Builders
If you are a founder or a developer, you shouldn't just look at this as 'another model release.' You should look at the implications of automated R&D. Here is how I see this impacting the ecosystem over the next 18 months:
- Accelerated Benchmarking: If we can automate the verification of new research, the gap between a discovery and its commercial application shrinks from years to months.
- Verification as a Service: There is a massive opportunity for startups to build 'trust layers' for AI-generated content and research. Faraday is the tool; the infrastructure around it is still up for grabs.
- The End of 'Paperware': We are entering an era where you can't just publish a theoretical paper and hope nobody notices the math doesn't work. Automated agents will be the new gatekeepers of scientific truth.
A Healthy Dose of Skepticism
Now, let's keep it real. Inherent is a startup, and startups need momentum. Claiming to beat Anthropic and OpenAI is the fastest way to get a meeting with a Tier-1 VC. While the benchmark results for Faraday look impressive, we need to see how it handles non-standardized data and papers that aren't already part of the common training sets.
The 'replication' task is a controlled environment. The real world is messy. Just because an agent can replicate a paper doesn't mean it can invent something new or navigate the political nuances of a corporate R&D department. We are seeing the 'what' and the 'how,' but the 'why' and the 'what next' are still very much in the hands of the humans running these labs.
The real breakthrough isn't an AI that can read a paper; it is an AI that can tell you why the paper is wrong before you waste six months trying to build on top of it.
The Founder Perspective
For those of us building in the crypto and AI overlap, this kind of tech is the missing link for decentralized science (DeSci). If you can have an autonomous agent verify research on-chain, you solve one of the biggest trust issues in the space. You don't need a central authority if you have a verifiable, reproducible agent that can prove the work was done correctly.
Inherent is playing a long game here. They aren't trying to replace the scientist; they are trying to replace the grunt work that prevents scientists from being creative. As a founder, that is the kind of leverage I am always looking for. I don't want an AI that writes my tweets; I want an AI that verifies my technical assumptions so I don't go down a three-month rabbit hole on a flawed premise.
Takeaway
Inherent's Faraday is a signal that the 'agentic' era is moving past the toy phase and into high-stakes environments like scientific research. While I'll wait for independent third-party verification before I declare them the new kings of the hill, the fact that a small lab can challenge the giants on specific, high-value tasks is a win for the entire builder community. It proves that specialized architecture still beats raw compute scale for complex reasoning.
Read the original at TechCrunch AI →