Skip to main content
AI Market Watch
Loading...

Datacurve

Category: AI Infrastructure

Datacurve provides expert-quality coding data for training and evaluating large language models (LLMs). Datacurve was founded in 2024. The company is led by Serena Ge. Based in San Francisco Bay Area, USA. Team size: 11-50. Total funding raised: $17.7M. Latest round: Series A. Key investors include Chemistry, Y Combinator, Afore Capital, Homebrew, Balaji Srinivasan, Northside Ventures, Palm Drive Capital, Pioneer Fund.

Founded
2024
Headquarters
San Francisco Bay Area, USA
Team size
11-50
Total funding
$17.7M

Value proposition

Provide high-quality, expert-vetted coding data that is difficult to source, enabling companies to build more capable and accurate AI models for software development.

Products and solutions

Shipd — gamified bounty platform sourcing expert coding/annotation data from skilled software engineers, DeepSWE — long-horizon coding agent benchmark (launched May 2026), RL environments, agent trajectories, OTS datasets, benchmarks/evals, and supervised fine-tuning data for frontier model labs

Unique value

Datacurve employs a bounty/gamified system to incentivize highly skilled software engineers and competitive programmers to produce expert-quality coding data, instead of relying on crowd-sourcing or synthetic data.

Target customer

Generative AI developer tool startups and other organizations building and fine-tuning LLMs for coding tasks.

Industries served

Artificial Intelligence, Software Development

Technology advantage

Superior data quality from expert contributors leads to more accurate and capable AI models; this differentiated dataset is the company's primary technical and business advantage.

How they differentiate

Focus on high-quality, expert-curated coding data using a bounty system to attract top developers for contribution and review.

Main competitors

Scale AI, Surge AI

Key partnerships

Y Combinator, Chemistry (lead investor in Series A)

Notable customers

Not publicly available

Major milestones

Backed by Y Combinator (W24), Raised $15M Series A led by Chemistry (Oct 2025); ~$17.7M total, Launched DeepSWE long-horizon coding agent benchmark (May 2026)

Growth metrics

Not publicly available

Market positioning

Positioned as a specialized provider of premium coding data, targeting the niche market of AI development focused on code generation and understanding, competing with larger generalist data platforms.

Geographic focus

Primarily North America, with a global reach.

Patents and IP

There is no publicly available information regarding patents or IP.

About Serena Ge

Co-Founder & CEO of Datacurve with a background in Computer Science from the University of Waterloo (dropped out to found Datacurve). Serena gained key industry experience as a Machine Learning Intern at Cohere, where she worked on training LLMs and reasoning capabilities, and through software engineering / related roles including Google (via co-founder collaboration context), with additional early experience building consumer products.

More AI Infrastructure companies

Official website: