Inherent, a London AI lab founded by former DeepMind researchers, emerged from stealth after a $50 million seed round and introduced its AI agent Faraday. The team says Faraday can independently reproduce the findings of published scientific papers without being given the answers, a capability that mirrors the early‑stage work of many PhD students.
In comparative tests, Faraday, built on the 27‑parameter‑billion Qwen 3.6 model, outperformed larger frontier systems such as Anthropic’s Claude Opus 4.8 and OpenAI’s GPT‑5.5 on this paper‑replication benchmark. Rather than relying on raw scale, the lab applied reinforcement learning that rewards successful experimental design, aiming to cultivate what the founders call “research taste.” To avoid duplicating existing tooling, Faraday leverages OpenAI’s GPT‑5.5 Codex for coding assistance.
- Model backbone: Qwen 3.6, 27 B parameters
- Compared systems: Claude Opus 4.8, OpenAI GPT‑5.5 (both larger)
- Training method: reinforcement learning focused on experimental‑design rewards
- Tool integration: uses GPT‑5.5 Codex for code generation
Why this matters
The outcome shows that a modest‑sized language model, when trained with reward‑based reinforcement learning targeting experimental design, can match or exceed the performance of substantially larger models on a narrow scientific‑replication task. This suggests that architectural efficiency and targeted training signals may be more decisive than parameter count for certain agentic capabilities. However, the claim is based on a single benchmark; broader generalization to novel hypothesis generation or cross‑domain discovery remains untested, so the result should be viewed as a promising proof‑of‑concept rather than evidence of universal superiority.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
