Sand.ai
Category: Foundation Models / LLMs
Sand.ai is a Chinese AI video generation company building autoregressive video models with audio-video sync, MoE architecture, and open-source releases, founded by Swin Transformer author Cao Yue. Sand.ai was founded in 2024. The company is led by Yue Cao (曹越). Based in Beijing, China. Team size: 11-50. Total funding raised: $100.0M. Latest round: Series B. Key investors include Look Capital, Lollapalooza Capital (Wang Huiwen Family Office), 九坤创投 (Jiukun Capital), 经纬创投 (Matrix Partners China), 和玉资本 (MSA Capital), 创新工场 (Sinovation Ventures), 襄禾资本, 源码资本 (Source Code Capital), 中科创星, 洪泰基金, 今日资本 (Today Capital), 华业天成.
- Founded
- 2024
- Headquarters
- Beijing, China
- Team size
- 11-50
- Total funding
- $100.0M
Value proposition
Open-source, SOTA video generation models using autoregressive architecture (not diffusion), with audio-video synchronization and MoE scaling for cost-efficient, high-quality video generation.
Products and solutions
Magi-1 (autoregressive video generation model, open-source), Gaga-1 (audio-video synchronized video model), VidMuse (music agent product, $10M ARR in 2 months), MagiAttention (distributed attention library for long-context training, open-source), MagiCompiler (optimization compiler for inference/training, open-source), Upcoming MoE video model (July 2026, open-source)
Unique value
Pioneer of autoregressive video generation (non-diffusion approach); first team outside Google to achieve audio-video synchronous output; Magi-1 ranked #1 on Google DeepMind's Physics IQ benchmark; open-source commitment with MagiAttention adopted by NVIDIA and most Chinese multimodal teams.
Target customer
AI video content creators, music video producers, enterprises needing video generation APIs, developers building video agents
Industries served
AI Video Generation, Music/Media Production, Digital Humans, Video Agents
Technology advantage
Autoregressive (AR) architecture for video (vs. mainstream diffusion); first-mover on audio-video synchronous generation; MoE architecture for video models (solving cost-speed-quality trilemma); MagiAttention library adopted by NVIDIA and most Chinese multimodal teams; open-source strategy driving ecosystem adoption; Swin Transformer lineage (ICCV 2021 Best Paper)
How they differentiate
Only major video generation company betting on autoregressive (next-frame prediction) rather than diffusion; first non-Google team to achieve audio-video synchronous output; strong open-source ethos with MagiAttention used by NVIDIA and Chinese multimodal teams; founder is Swin Transformer author (50K+ citations, ICCV Best Paper); lean team (<30) with high research productivity
Main competitors
ByteDance Seedance, Kuaishou Kling, OpenAI Sora, Google Veo
Key partnerships
NVIDIA (MagiAttention recommended for multimodal model training), Open-source community (HuggingFace, GitHub), 星涵资本 (Xinghan Capital) as financial advisor for B round
Major milestones
2024-01: Company founded by Cao Yue, 2024-07: Seed round from Source Code Capital & Today Capital, 2025-04: Released Magi-1 (first high-quality autoregressive video model, open-source), 2025: Magi-1 ranked #1 on Google DeepMind Physics IQ benchmark, 2025-11: Decision to pivot from Dense to MoE architecture, 2026-01: Launched VidMuse music agent product, 2026-03: VidMuse reached $10M ARR in 2 months, 2026-04: $50M A++ round, 2026-06: B round closing >$100M total across two rounds, 2026-Q3: Planned release of MoE video model (open-source)
Growth metrics
VidMuse achieved $10M ARR within 2 months of launch (Jan 2026); Magi-1 open-source repo has 3.7K+ GitHub stars; MagiAttention has 859+ GitHub stars
Market positioning
Leading Chinese open-source AI video generation company competing with ByteDance (Seedance), Kuaishou (Kling), and global players (OpenAI Sora, Google Veo) through autoregressive architecture, audio-video sync, and MoE scaling.
Geographic focus
China (primary), Global (open-source distribution)
About Yue Cao (曹越)
Co-founder of Lightyear AI (光年之外, acquired by Meituan 2023); Head of Multimodal & Vision Research Center at BAAI (Beijing Academy of AI); Senior Researcher at Microsoft Research Asia (2019-2022); PhD from Tsinghua University (2019); Author of Swin Transformer (ICCV 2021 Best Paper / Marr Prize)
Latest news about Sand.ai
More Foundation Models / LLMs companies
Official website: https://sand.ai