
Sapient Intelligence reports 49 MLE-bench golds for Praxist at lower model cost
The AMW Read
The reported benchmark meaningfully updates the research-agent player map, but its broader impact depends on isolating the system's contribution from model choice and validating results beyond this test.
Sapient Intelligence reports 49 MLE-bench golds for Praxist at lower model cost
Singapore-based Sapient Intelligence says Praxist, its autonomous research system, earned 49 gold medals across OpenAI's 75-task MLE-bench at about $3,054 in model costs. A Claude Code baseline using Claude Opus 4.8 earned 34 golds at $38,370, according to the company's release and technical paper. Praxist ran on DeepSeek-v4-pro, so the comparison changed both the research system and the underlying model. The paper also says an internal integrity review set aside 90,423 experiments and replaced the best score on nine tasks with a verified score, costing four medals.
Praxist belongs to the emerging market for agents that pursue open-ended goals through repeated actions and feedback. Its proposed distinction is a research graph that retains failed experiments and combines lessons across attempts, rather than pruning weaker branches. That could matter for data science work where the best approach is unknown at the start. The benchmark, however, does not isolate the graph's contribution: co-founder William Chen told CDOTrends he could not quantify its share of the result. The reported cost gap likewise cannot establish that Praxist's design, rather than model choice, caused the savings.
Builders evaluating research agents should ask for a same-model comparison, an ablation of the research graph, and an auditable record from initial experiment through final score. Investors should distinguish model spend on this benchmark from the broader cost of operating the system. Those checks would show whether Praxist's experiment memory offers a repeatable advantage beyond this reported result.