XtalPi Claims Its Science Agent Platform Beats Claude on Drug Discovery Benchmarks
The AMW Read
XtalPi's self-reported 90%-vs-70% benchmark against Claude and new robotic-lab dispatch extend the emerging agentic-science-lab pattern Anthropic opened with Claude Science, updating segment 06's competitive baseline without resolving any open debate.
XtalPi Claims Its Science Agent Platform Beats Claude on Drug Discovery Benchmarks
XtalPi (晶泰科技), the Shenzhen-based, Hong Kong-listed (2228.HK) drug discovery firm, upgraded its science agent platform, XtalPi Science, adding 80+ proprietary skills and tools and widening invite-only access. In an expert-blind-reviewed hit-to-lead workflow test, it hit a 90% "effective recommendation" rate versus 70% for Anthropic's Claude, and a 60% "excellent" rate versus roughly 20%. In a separate synthesis-route test, it proposed a regioselective route where a comparison general-purpose agent misjudged the starting material. The platform now also dispatches robots in XtalPi's own labs, running experiments and adjusting tasks from results; on one internal project this cut human analysis workload 80%, letting a 15-person chemist team deliver over 20,000 reaction records monthly.
The claim sits inside a fast-forming race to turn foundation models into working lab scientists. Anthropic launched Claude Science in late June 2026, bundling research tools, data, and compute, and has since built wet labs for Claude to direct robotic experiments; Google DeepMind's AI co-scientist pursues a similar multi-agent approach to hypothesis generation. XtalPi's bet is that proprietary drug-discovery data and packaged expert workflows let a narrower, domain-trained agent beat a general frontier model on real pharma tasks — per the AI Market Watch index, XtalPi logged just one prior pipeline-matched story in the last 90 days, against zero before, making this a fresh line of coverage.
For builders and investors, the signal is that proprietary data, codified expert workflows, and owned lab robotics are becoming differentiators against general-purpose agents in execution-heavy, regulated verticals like drug discovery. The benchmark is self-reported and internal, so treat the win-rate numbers as a competitive claim pending independent validation.