
Inherent's Faraday agent outperforms frontier models at replicating scientific research
The AMW Read
A DeepMind-alumni science-agent startup shows RL-trained small models can outperform larger frontier systems on a specific benchmark, meaningfully updating the model-size-vs-training-method debate without resolving it.
Named counterparties: OpenAI
Inherent's Faraday agent outperforms frontier models at replicating scientific research
London-based Inherent, founded by Google DeepMind alumni including chief scientist Edward Hughes, said its AI agent Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at independently reproducing the findings of published scientific papers without prior access to the results. The task, a standard exercise for early-career researchers, tests not just accuracy but what Inherent calls "research taste" β judgment about which experiments are worth running and how to design them. Faraday runs on a 27-billion-parameter Qwen 3.6 model, far smaller than the frontier systems it beat, trained with reinforcement learning rather than imitation of how science is described in text. The startup emerged from stealth weeks earlier with a $50 million seed round and employs about a dozen people out of a King's Cross office, with plans to grow to 20-25 by year-end. Per the AI Market Watch index, name-matched pipeline coverage of Inherent rose to 7 items in the past 90 days from 5 in the prior 90, over pipeline-ingested sources only.
The result complicates the assumption that frontier-scale general models are the default choice for specialized reasoning tasks. Inherent's bet is that reward-based training aimed at judgment, not parameter count, generalizes better toward autonomous scientific discovery, and that a lean RL-trained agent built on an open model can beat systems with far larger training budgets on a narrowly defined benchmark. The company also declined to build its own coding tool, routing Faraday through OpenAI's GPT-5.5 Codex instead, treating training methodology as the differentiator rather than owning every layer of the stack.
For investors and builders, the signal is that model size is becoming a weaker predictor of task performance than training method and evaluation design on narrow, well-specified benchmarks β worth testing before assuming frontier-model subscriptions are the only credible path to research-grade agents. London's DeepMind-trained talent pool, and staff reportedly unsettled by Demis Hassabis's new role, also make Inherent a hiring destination larger labs should watch.