
OpenArt Launches Use-Case Benchmark Arena for Image and Video Generation Models
The AMW Read
A new per-task evaluation methodology for image/video generation reshuffles competitive framing among ByteDance, Alibaba, and OpenAI without a new top-tier entrant or debate resolution.
OpenArt Launches Use-Case Benchmark Arena for Image and Video Generation Models
OpenArt, a creative AI platform founded by Google alums Coco Mao and John Cao that aggregates third-party image, video, voice, and 3D generation models, launched OpenArt Arena on September 15. Rather than one aggregate leaderboard, it ranks models by production task — film, e-commerce, graphic design, motion design, video editing, lip-sync — using blind pairwise comparisons scored with the Bradley-Terry method and judged by 800 to 1,000 vetted creative professionals and outside evaluators. In video, ByteDance's Seedance 2.5 led overall at 1,081 points and topped film, motion design, and lip-sync, though Alibaba's Wan 3.0 narrowly beat it in video editing, 1,034 to 1,033. In image, OpenAI's GPT-Image-2 led graphic design and image editing, while Seedream 5.0 Pro topped film-style and e-commerce imagery and edged GPT-Image-2 for the overall image lead, 1,010 to 1,000.
Per-task rather than aggregate scoring is the notable shift: a model built for cinematic camera movement is judged on different criteria than one built for legible ad copy or accurate product placement, and no single model won across categories. OpenArt discloses only part of its prompt set — holding the rest private to stop developers from optimizing against the public test — but hasn't published total prompt counts or vote volume, leaving transparency and reproducibility open questions. Because OpenArt hosts the models it ranks on its own platform, the arena also isn't fully independent of the vendors being scored.
For builders, the result argues against defaulting to one model across a production pipeline — the winner changed by deliverable, reinforcing per-task model routing over vendor loyalty. For OpenArt and platforms like it, the harder question is whether they can stay a credible, neutral arbiter while also selling access to the same models they're ranking.
