Skip to main content
Back to News
Perplexity AI launched DRACO, an open-source benchmark evaluating research agents via 100 tasks from...
Technology
1 min read
US

Perplexity AI launched DRACO, an open-source benchmark evaluating research agents via 100 tasks from...

The AMW Read

Perplexity (a key player in research/search) is updating the agentic evaluation landscape with a production-grounded benchmark, but this is an incremental tool release rather than a structural shift.
NoveltySignificance
AI Agents · Player Map

Perplexity AI launched DRACO, an open-source benchmark evaluating research agents via 100 tasks from real user queries. Spanning 10 domains, Perplexity leads with 89.4 percent accuracy in Law and 82.4 percent in Academic research. Shifting from synthetic puzzles to production-grounded data creates a rigorous standard for multi-step reasoning. This systemic evolution forces the AI industry to prioritize factual depth over conversational fluency. 🚀

#AIResearch #DRACO #PerplexityAI #LLM #Technology

How This Connects

Based on AI Agents · Player Map

  1. 4d agoEveryone’s AI brings Korea’s telecom and messaging platforms into consumer agentsEveryone’s AI
  2. 1w agoManus launches Manus 2.0 and Cue personal-agent appManus
  3. 1w agoInstinct reportedly raises $1B Series C at $10B valuationInstinct
  4. 1w agoNvidia Expands OpenShell and Introduces Sentry for AI Agent SecurityNvidia
  5. 1mo agoMCP's largest spec revision since launch ships with a prompt-injection flaw that can leak credentials.Model Context Protocol (MCP) spec update
  6. 8mo agoPerplexity AI launched DRACO, an open-source benchmark evaluating research agents via 100 tasks from... · THIS ARTICLE

Related News

More news from Perplexity

Stay updated with the latest news and announcements from Perplexity.

View all Perplexity news

Discover AI Startups

Explore 5,000+ AI companies with VC-grade analysis, funding data, and investment insights.

Explore Dashboard