HiDream ships HiDream-O1-Video-1.0 with dual top-tier image-to-video board rankings
The AMW Read
Dual I2V board debut completes HiDream’s O1 multimodal matrix and updates the generative-video player map; C+ lacks a disclosed size so no capital cross-tag.
HiDream ships HiDream-O1-Video-1.0 with dual top-tier image-to-video board rankings
HiDream.ai on September 15 released HiDream-O1-Video-1.0 (HD-V1), its first native full-modal video generation model under the HiDream-O1 stack. HD-V1 takes text, image, and video inputs and can output 5–20 second 1080p clips, with joint modeling of text, video, and audio rather than post-hoc soundtrack splicing. On Artificial Analysis’s Image-to-Video leaderboard (With Audio) the model ranked fourth globally; on Arena Image-to-Video it ranked eighth. Alongside the launch, HiDream said it closed a Series C+ round from Xinwei Capital, Jiaozi Capital, and ICBC Capital; the round size was not disclosed. Per the AI Market Watch index, HiDream.ai is a 2023-founded lab with about $83.0M in total funding tracked across the index’s ~5,000 companies.
The release matters because image-to-video remains one of the densest competitive surfaces in generative media, and dual independent-board placement on first public outing is how new Chinese labs force Western and domestic peers to recalibrate who sits in the first tier. HD-V1 also extends HiDream’s recent O1 push: after HiDream-O1-Embodied topped a RoboColiseum robustness board earlier this month, Video fills the missing piece of the firm’s stated image–video–interactive–embodied matrix under one native multimodal architecture. Claims around intent planning, physics-aware motion, content-driven duration instead of fixed clip windows, and Diffusion Reinforcement Learning with a multimodal reward model are the differentiators buyers will pressure-test against Runway-class and domestic video rivals.
For builders and investors, the near-term signal is whether native audio-visual joint generation plus content-tied duration become procurement requirements rather than demos. Capital already appears willing to underwrite the “one native multimodal base, four co-evolving models” thesis even without a published C+ figure; the next proof point is whether HD-V1’s leaderboard position holds once rivals refresh and whether commercial API or enterprise workflows follow the benchmark splash.