TokenRhythm releases NeoHorse-1, a 4B/9B model trained on agent-execution trajectories from its Routing Harness system
The AMW Read
A non-top-tier lab introduces trajectory-curated post-training plus an explicit training-scaling-curve claim on open-weight Qwen3.5, a real if small-scale technical proof point rather than a debate-resolving event.
TokenRhythm releases NeoHorse-1, a 4B/9B model trained on agent-execution trajectories from its Routing Harness system
TokenRhythm (基元律动), founded by former Huawei Noah's Ark Lab director and Pangu large-model lead Wang Yunhe, released NeoHorse-1 with partners Infinigence AI, Tsinghua University, Peking University and Alibaba. The 4B and 9B versions are post-trained on Alibaba's open-source Qwen3.5-4B and Qwen3.5-9B bases. Training data centers on execution trajectories from TokenRhythm's open-source Routing Harness system, OpenSquilla, which selects and orchestrates models as agents work — trajectories retain capability-requirement predictions, routing choices, model responses, tool calls and environment feedback, filtered through completeness, goal-completion, evidence-consistency and error-recovery checks, then used for post-training via on-policy distillation. Across 11 benchmarks covering agent execution, tool use, coding and instruction-following, the 4B model posted the highest unweighted average among 4B-class peers and beat the Qwen3.5-9B base on five benchmarks, with gains concentrated in tasks with clear process and verifiable outcomes; larger models still lead on complex state management and long-horizon debugging.
The release ties agent execution data back into model training in what TokenRhythm frames as a single-round validation of recursive self-improvement: agents act, evaluation exposes capability gaps, training selection adjusts, and the updated model returns to the harness. It treats agent operating experience as training signal for the base model rather than routing being a thin layer atop a static model, using Alibaba's open Qwen3.5 stack and academic partners to test whether that loop compounds into a steeper capability-per-training-investment curve.
For teams evaluating cost-sensitive agent models, NeoHorse suggests trajectory-quality curation can narrow gaps with larger bases on well-instrumented, verifiable tasks, though not yet on long-horizon failure recovery. Per the AI Market Watch index, our pipeline has logged two TokenRhythm items in the past 90 days, up from zero prior (name-matched, pipeline-ingested sources only) — an August round funding multi-model routing infrastructure, now followed by a first technical proof that routing/orchestration data can feed back into model quality, the bet investors funded.
#TokenRhythm #NeoHorse #AgentNativeAI #Qwen #OpenSource #RecursiveSelfImprovement