
StepFun releases Step 5 Preview, claims top-3 global open-source rank at a fraction of Claude Opus 5's cost
The AMW Read
A flagship MoE model claiming top-3 open-source rank and roughly one-eighth Claude Opus 5's per-task cost meaningfully updates StepFun's competitive position and its deliberate open-weight release strategy, without resolving the broader frontier-vs-open debate.
StepFun releases Step 5 Preview, claims top-3 global open-source rank at a fraction of Claude Opus 5's cost
StepFun (阶跃星辰) released Step 5 Preview on September 20, a new flagship base model built for real-world agentic work — coding, software engineering, professional knowledge tasks, and finance. The model uses a sparse mixture-of-experts design with 600 billion total parameters and only 27 billion active per token, a 1-million-token context window, and native text-and-vision input. Full weights are scheduled to open source on October 15.
On the Artificial Analysis Intelligence Index, Step 5 Preview scored 44, placing it among the top three open-source models globally, while StepFun says its per-task inference cost runs roughly one-eighth of Claude Opus 5's. On agentic and coding benchmarks — DeepSWE v1.1 (67.7%) and StepCodeBench (49.0%) — plus the CLI subset of Agents' Last Exam and a financial-research benchmark, the model trails only GPT-6 Astra or Claude Opus 5 among evaluated systems, ahead of other open-weight entrants. Per the AI Market Watch index, StepFun has raised $3,800.0M in total funding to date — a figure from our index's coverage, not a full census, but it signals the capital scale behind this release.
For builders, an open-weight model claiming near-frontier agentic and finance benchmark results at a fraction of closed-model cost widens the option set for cost-sensitive coding-agent and back-office deployments once the weights ship in October. For investors, Step 5 Preview extends StepFun's push beyond single-modality releases — following its recent voice-model and on-device model launches — into a broader model stack, making per-task cost, not just leaderboard rank, the metric to watch as open-weight labs compete directly with frontier closed models on agentic workloads.
