Skip to main content

LatentVerse

Category: Robotics / Embodied AI

Beijing-based embodied AI foundation model startup developing UTAM (Unified Tactile Action Model), a unified architecture integrating vision-language understanding, world modeling, action generation, and tactile feedback for robots. LatentVerse was founded in 2026. The company is led by Hu Yucheng (胡钰承). Based in Beijing, China. Team size: 11-50. Total funding raised: $50.0M. Latest round: Seed. Key investors include GL Ventures (高瓴创投), Qingliu Capital (清流资本), InnoAngel Fund (英诺天使基金), Zhiyuan/Agibot (智元机器人), StarMotion Era/Robot Era (星动纪元).

Founded
2026
Headquarters
Beijing, China
Team size
11-50
Total funding
$50.0M

Value proposition

Builds an "embodiment-native" foundation model (UTAM) that unifies VLM, world model, action, and tactile sensing in a single architecture — a third path beyond pure VLA and pure world-model (WAM) approaches. Combines cognitive intelligence (understanding/planning) with contact intelligence (real-time tactile closed-loop correction) to enable generalization, long-horizon execution, and dexterity simultaneously.

Products and solutions

UTAM (Unified Tactile Action Model) — embodied full-modality foundation model evolved from the team's earlier UAM (Unified Action Model), outputs visual, language, and tactile feedback signals simultaneously. Prior research lineage: DP3, HiRT, VPP, BagelVLA, UAM.

Unique value

Unified Tactile Action Model (UTAM) with four expert modules (vision-language, world model, action, high-frequency tactile) that fuses open-loop execution with closed-loop tactile correction — solving the "last centimeter" contact problem that pure vision-based VLA/world models cannot handle.

Target customer

Embodied AI / robotics companies and robot OEMs (e.g., investors Zhiyuan/Agibot and StarMotion Era/Robot Era are industry players); target scenarios include hotel room cleaning, packing in cramped spaces, home grasping of soft objects.

Industries served

Embodied AI / robotics foundation models, home and service robotics, industrial manipulation

Technology advantage

Unified architecture that can train on a far larger data pool than pure VLA (manipulation videos, cross-embodiment robot data, general VLA data, internet-scale video). Tactile expert enables high-frequency closed-loop contact correction (force, slippage, object deformation) that vision alone cannot capture. Team includes researchers who pioneered WAM (VPP) and led BagelVLA at ByteDance Seed.

How they differentiate

Unlike VLA (data-hungry, action-annotation dependent) and pure world models (lack semantic understanding), LatentVerse unifies VLM + world model + action + tactile in one architecture. Its tactile expert provides high-frequency closed-loop contact correction, addressing the "last centimeter" dexterity problem that vision-only systems cannot solve.

Main competitors

Physical Intelligence (Pi-0.5, VLA route), NVIDIA (DreamZero, WAM route), Figure (Helix dual-system architecture), other Chinese embodied foundation model startups (e.g., Galaxy General/银河通用)

Key partnerships

Strategic industry investors Zhiyuan/Agibot (智元机器人) and StarMotion Era/Robot Era (星动纪元), team talent from ByteDance Seed, Alibaba Qwen, Xiaomi, Tsinghua IIIS, Peking University, NTU

Major milestones

Founded May 2026 by Tsinghua IIIS embodied AI research team, completed several-hundred-million-RMB seed round led by GL Ventures (announced Aug 11, 2026), began joint training of UTAM expert modules with multimodal data infrastructure and training pipelines built.

Market positioning

Early-stage (seed) embodied AI foundation model startup taking a differentiated "third path" beyond VLA and world models, emphasizing tactile/contact intelligence. Backed by top-tier investors (GL Ventures/Hillhouse) and strategic industry players (Zhiyuan/Agibot, StarMotion Era/Robot Era).

Geographic focus

China (Beijing), with global relevance in embodied AI foundation models

About Hu Yucheng (胡钰承)

PhD candidate at Tsinghua University's Institute for Interdisciplinary Information Sciences (IIIS), mentored by Chen Jianyu (founder of Robot Era/星动纪元); ex-ByteDance Seed where he led R&D of the BagelVLA series of embodied foundation models. Born 2001. Proposed PAD (multimodal diffusion VLA, early 2024) and VPP (first video-action model with closed-loop reasoning, Nov 2024), the latter considered pioneering work of the World Action Model (WAM) route later adopted by NVIDIA's DreamZero.

Latest news about LatentVerse

More Robotics / Embodied AI companies