Datacurve
Category: AI Infrastructure
Datacurve provides expert-quality coding data for training and evaluating large language models (LLMs). Datacurve was founded in 2024. The company is led by Serena Ge. Based in San Francisco Bay Area, USA. Team size: 11-50. Total funding raised: $17.7M. Latest round: Series A. Key investors include Chemistry, Y Combinator, Afore Capital, Homebrew, Balaji Srinivasan, Northside Ventures, Palm Drive Capital, Pioneer Fund.
- Founded
- 2024
- Headquarters
- San Francisco Bay Area, USA
- Team size
- 11-50
- Total funding
- $17.7M
Value proposition
Provide high-quality, expert-vetted coding data that is difficult to source, enabling companies to build more capable and accurate AI models for software development.
Products and solutions
Shipd — gamified bounty platform sourcing expert coding/annotation data from skilled software engineers, DeepSWE — long-horizon coding agent benchmark (launched May 2026), RL environments, agent trajectories, OTS datasets, benchmarks/evals, and supervised fine-tuning data for frontier model labs
Unique value
Datacurve employs a bounty/gamified system to incentivize highly skilled software engineers and competitive programmers to produce expert-quality coding data, instead of relying on crowd-sourcing or synthetic data.
Target customer
Generative AI developer tool startups and other organizations building and fine-tuning LLMs for coding tasks.
Industries served
Artificial Intelligence, Software Development
Technology advantage
Superior data quality from expert contributors leads to more accurate and capable AI models; this differentiated dataset is the company's primary technical and business advantage.
How they differentiate
Focus on high-quality, expert-curated coding data using a bounty system to attract top developers for contribution and review.
Main competitors
Scale AI, Surge AI
Key partnerships
Y Combinator, Chemistry (lead investor in Series A)
Notable customers
Not publicly available
Major milestones
Backed by Y Combinator (W24), Raised $15M Series A led by Chemistry (Oct 2025); ~$17.7M total, Launched DeepSWE long-horizon coding agent benchmark (May 2026)
Growth metrics
Not publicly available
Market positioning
Positioned as a specialized provider of premium coding data, targeting the niche market of AI development focused on code generation and understanding, competing with larger generalist data platforms.
Geographic focus
Primarily North America, with a global reach.
Patents and IP
There is no publicly available information regarding patents or IP.
About Serena Ge
Co-Founder & CEO of Datacurve with a background in Computer Science from the University of Waterloo (dropped out to found Datacurve). Serena gained key industry experience as a Machine Learning Intern at Cohere, where she worked on training LLMs and reasoning capabilities, and through software engineering / related roles including Google (via co-founder collaboration context), with additional early experience building consumer products.
More AI Infrastructure companies
Official website: https://datacurve.ai/