LatentVerse
Category: Robotics / Embodied AI
Beijing-based embodied AI foundation model startup developing UTAM (Unified Tactile Action Model), a unified architecture integrating vision-language understanding, world modeling, action generation, and tactile feedback for robots. LatentVerse was founded in 2026. The company is led by Hu Yucheng (胡钰承). Based in Beijing, China. Team size: 11-50. Total funding raised: $50.0M. Latest round: Seed. Key investors include GL Ventures (高瓴创投), Qingliu Capital (清流资本), InnoAngel Fund (英诺天使基金), Zhiyuan/Agibot (智元机器人), StarMotion Era/Robot Era (星动纪元).
- Founded
- 2026
- Headquarters
- Beijing, China
- Team size
- 11-50
- Total funding
- $50.0M
Value proposition
Builds an "embodiment-native" foundation model (UTAM) that unifies VLM, world model, action, and tactile sensing in a single architecture — a third path beyond pure VLA and pure world-model (WAM) approaches. Combines cognitive intelligence (understanding/planning) with contact intelligence (real-time tactile closed-loop correction) to enable generalization, long-horizon execution, and dexterity simultaneously.
Products and solutions
UTAM (Unified Tactile Action Model) — embodied full-modality foundation model evolved from the team's earlier UAM (Unified Action Model), outputs visual, language, and tactile feedback signals simultaneously. Prior research lineage: DP3, HiRT, VPP, BagelVLA, UAM.
Unique value
Unified Tactile Action Model (UTAM) with four expert modules (vision-language, world model, action, high-frequency tactile) that fuses open-loop execution with closed-loop tactile correction — solving the "last centimeter" contact problem that pure vision-based VLA/world models cannot handle.
Target customer
Embodied AI / robotics companies and robot OEMs (e.g., investors Zhiyuan/Agibot and StarMotion Era/Robot Era are industry players); target scenarios include hotel room cleaning, packing in cramped spaces, home grasping of soft objects.
Industries served
Embodied AI / robotics foundation models, home and service robotics, industrial manipulation
Technology advantage
Unified architecture that can train on a far larger data pool than pure VLA (manipulation videos, cross-embodiment robot data, general VLA data, internet-scale video). Tactile expert enables high-frequency closed-loop contact correction (force, slippage, object deformation) that vision alone cannot capture. Team includes researchers who pioneered WAM (VPP) and led BagelVLA at ByteDance Seed.
How they differentiate
Unlike VLA (data-hungry, action-annotation dependent) and pure world models (lack semantic understanding), LatentVerse unifies VLM + world model + action + tactile in one architecture. Its tactile expert provides high-frequency closed-loop contact correction, addressing the "last centimeter" dexterity problem that vision-only systems cannot solve.
Main competitors
Physical Intelligence (Pi-0.5, VLA route), NVIDIA (DreamZero, WAM route), Figure (Helix dual-system architecture), other Chinese embodied foundation model startups (e.g., Galaxy General/银河通用)
Key partnerships
Strategic industry investors Zhiyuan/Agibot (智元机器人) and StarMotion Era/Robot Era (星动纪元), team talent from ByteDance Seed, Alibaba Qwen, Xiaomi, Tsinghua IIIS, Peking University, NTU
Major milestones
Founded May 2026 by Tsinghua IIIS embodied AI research team, completed several-hundred-million-RMB seed round led by GL Ventures (announced Aug 11, 2026), began joint training of UTAM expert modules with multimodal data infrastructure and training pipelines built.
Market positioning
Early-stage (seed) embodied AI foundation model startup taking a differentiated "third path" beyond VLA and world models, emphasizing tactile/contact intelligence. Backed by top-tier investors (GL Ventures/Hillhouse) and strategic industry players (Zhiyuan/Agibot, StarMotion Era/Robot Era).
Geographic focus
China (Beijing), with global relevance in embodied AI foundation models
About Hu Yucheng (胡钰承)
PhD candidate at Tsinghua University's Institute for Interdisciplinary Information Sciences (IIIS), mentored by Chen Jianyu (founder of Robot Era/星动纪元); ex-ByteDance Seed where he led R&D of the BagelVLA series of embodied foundation models. Born 2001. Proposed PAD (multimodal diffusion VLA, early 2024) and VPP (first video-action model with closed-loop reasoning, Nov 2024), the latter considered pioneering work of the World Action Model (WAM) route later adopted by NVIDIA's DreamZero.