
OpenAI Proposes Four-Metric 'Useful Intelligence per Dollar' Framework to Measure Enterprise AI ROI
The AMW Read
OpenAI introduces a structured four-metric ROI framework and discloses GPT-5.6 tiered pricing and benchmark data, meaningfully advancing the segment's shift from token-based to outcome-based value framing for foundation-model buyers.
OpenAI Proposes Four-Metric 'Useful Intelligence per Dollar' Framework to Measure Enterprise AI ROI
OpenAI on July 17, 2026 published a framework it calls Useful Intelligence per Dollar, arguing that pricing AI purely by token consumption understates real cost efficiency. The framework scores AI deployments on four metrics: business outcomes achieved, total cost incurred, reliability, and value delivered as usage scales. It ties into GPT-5.6, released July 9, 2026 in three tiers — Sol, Terra, and Luna — with Luna priced roughly 80% below Sol and Terra roughly 20% below, aimed at high-volume, cost-sensitive tasks. On the Artificial Analysis Coding Agent Index, GPT-5.6 Sol set a new record at maximum reasoning settings, scoring 72.7 against DeepSWE v1.1 while using 36.2% fewer tokens via API than Anthropic's Claude Fable 5; company-wide, OpenAI's output token volume fell 54%.
The framework gives OpenAI a vocabulary to compete on operational cost rather than headline capability, distributed through ChatGPT Work, which layers access controls onto internal business systems for enterprise deployment. It lands as OpenAI has been narrowing Anthropic's lead in U.S. business AI spending — Ramp data cited in recent coverage put OpenAI at roughly 44% share versus Anthropic's ~40% in July, after pushback over Anthropic's Fable-tier pricing and data retention. Tiered pricing across Sol, Terra, and Luna lets OpenAI segment budget-constrained accounts from those paying for peak reasoning, while the four-metric framing reframes procurement conversations around outcomes instead of per-token rates.
For enterprise buyers, the framework is a negotiating tool: it pressures vendors to disclose reliability and correction-cost figures rather than raw benchmark scores, and gives procurement teams a basis to compare OpenAI against Anthropic and other providers on total cost per outcome rather than list price. For competitors, the risk is definitional — if enterprises adopt OpenAI's four metrics as a standard, OpenAI controls the terms its own products are judged by.

