
Snorkel AI closes $350M Series E at $3.5B valuation as training-data revenue scales 18x.
The AMW Read
Valuation nearly triples in 17 months and revenue grows 18x, and the article surfaces a structural gross-vs-net revenue distinction that separates Snorkel's dataset-sale model from marketplace-style AI-data-labor peers.
Snorkel AI closes $350M Series E at $3.5B valuation as training-data revenue scales 18x.
Snorkel AI, a seven-year-old startup that builds training datasets and simulated environments for AI labs and enterprises, closed a $350M Series E led by Insight Partners and S32 at a $3.5B valuation. That is nearly triple the $1.3B valuation the company held 17 months ago at its Series D, when it raised $100M. Existing investors Addition, Lightspeed, Greylock, GV, and Wells Fargo also participated. Snorkel says its annualized revenue run rate has reached $375M, up 18x over the past 12 months, which it attributes to strong demand from AI labs for high-end training data.
Snorkel originally sold data-labeling automation software, but pivoted last year to what it calls a data-as-a-service model, selling complete datasets and reinforcement-learning environments built through a hybrid of its own software and models plus domain experts, rather than operating as a pure human-labor marketplace. That structural choice lets Snorkel book payments to its domain experts as cost of goods sold rather than netting them against revenue. The distinction matters because peers in the same AI-data-lab category report striking but structurally different numbers: Mercor's gross annualized revenue has climbed to $2B, Handshake crossed $1B earlier this year, and Micro1 has reportedly scaled to $500M, yet companies in this category typically pass 60-70% of revenue straight through to the experts doing the work, making net revenue far lower than the gross totals suggest.
For investors comparing companies in this category, gross annualized revenue is not a reliable cross-company metric when labor-cost structures differ this much; Snorkel's bet is that selling finished datasets and RL environments outright, rather than brokering labor, can preserve better margins as frontier labs keep paying up for high-end training data.