
Apto launches A/B testing service to verify whether costly AI models justify their price
The AMW Read
A small pre-launch student startup adds a live A/B cost-vs-outcome measurement layer to LLM observability tooling, a new entrant on the player map but too early-stage to shift the segment baseline.
Apto launches A/B testing service to verify whether costly AI models justify their price
South Korean startup Apto (앱토), founded by three university students including Korea University computer science student Kang Dong-hyuk, launched a service on September 23 that helps companies test whether upgrading to a more expensive AI model actually pays off. Rather than relying on the coding, math, and knowledge benchmark scores that model developers publish, Apto splits real users of the same AI-powered feature into groups running different models or prompts, then measures outcomes such as payment conversion and return visits against the AI inference cost each group generates. A company can roll a new model out to just 10 percent of users first: if results don't improve, it can shift more traffic to a cheaper model, and if the incumbent model performs meaningfully better, that gap becomes the evidence for keeping the higher spend.
The service targets a widening gap between AI spend and measured return. Apto cites a Menlo Ventures report showing US enterprise generative AI spending rose from $1.7 billion in 2023 to $37 billion in 2025, and an MIT Media Lab NANDA project report finding that 95 percent of organizations saw no bottom-line return from their generative AI deployments. That gap has turned model selection into a live measurement problem rather than a one-time benchmark comparison: as frontier labs ship new model tiers every few months, the benchmark score gap between a flagship and a lighter model tells buyers little about whether a specific feature and user base actually need the more expensive option.
For enterprise buyers, this points to model-routing and evaluation tooling as a companion layer to the LLM stack rather than a one-off vendor comparison. For investors, Apto is not yet a signal on its own: it is a pre-launch team out of Korea's government-run AI-SW Maestro training program with no disclosed funding or customers, so the more durable read is the demand pattern it is responding to, not the company itself.