Home/AI Infrastructure/ShengShu Technology
Company

ShengShu Technology

Beijing "general world model" startup — Vidu video generation plus Motus/Motus2 dexterous-manipulation world models; Alibaba-backed unicorn [5].

1. Core Product / Service

ShengShu Technology (生数科技, founded March 2023) builds what it calls a general world model, spanning "multimodal generation + real-time interaction + physical-world action" [5]. Its thesis: high-quality video generation already encodes an understanding of physics, so the same foundation extends from the digital world into the physical one [5].

Vidu is the video-generation line — a series of models with synchronized audio-video generation, long-duration output and high spatiotemporal consistency, serving text/image-to-video and reference-material-to-video through a MaaS open platform plus SaaS tools (Vidu Agent, Vidu Claw). By March 2025 Vidu had users in 200+ countries/regions; the Vidu Q series runs across manhua, short drama, ads, e-commerce, animation and film pipelines [5].

Motus / MotuBrain / Motus2 extend this into embodied AI. Motus (December 2025, open-sourced) is billed as the first unified "perception–prediction–action" latent-action world model — usable both as a world-action model (generating actions) and an action-conditioned world model (predicting the visual state after an action) [5]. MotuBrain (April 2026) is the commercial follow-on that ShengShu "claimed" after it quietly topped two benchmarks: #1 across motion-quality and smoothness dimensions on WorldArena, and 95.8% (clean) / 96.1% (randomized) average success on RoboTwin 2.0 — the only model averaging above 95 in randomized environments [6]. Motus2 (announced September 10, 2026, at the Bund Conference) is a "self-evolving general world model for dexterous manipulation," built with Tsinghua University and tested on dual-arm robots using WUJI Hand 2 and Sharpa Wave dexterous hands [7].

Motus2 folds policy generation, world simulation (consequence prediction) and value assessment into one closed loop: it proposes actions, predicts outcomes, and evaluates which future is closer to the goal, then uses that feedback to self-improve. It extracts manipulation priors from ~130,000 hours of human first-person video and adapts them to robot embodiments with trajectory data; on the execution side it adds binocular vision, working memory and tactile sensing, plus a value model and an "action-first" causal information flow [7].

2. Target Users & Pain Points

  • Robot / embodied-AI developers who need training environments and world models for dexterous manipulation — the hardest part of robotics — rather than building simulators from scratch.
  • Content producers (ads, film, short drama, e-commerce) using Vidu for commercial video generation.

Pain solved: the sim2real and data-scarcity bottleneck in manipulation. Jimmy's late-September research frames the field as locomotion far more mature than manipulation, with tight-sim2real world models the lever to close that gap [2][3].

3. Competitive Landscape

Player Focus Vs. ShengShu
ShengShu Video gen (Vidu) + manipulation world models (Motus2) Video→world-model continuum; open-sourced Motus
galaxea Embodied VLA (G0.5) + hardware Unified think-while-act VLA; sells robots
physical-intelligence Unified VLA (π series) Closed-source; robotics-native (no video-gen line)
minimax Hailuo video + H3 world model Broad multimodal lab; H3 as long-sequence action sim

ShengShu's differentiation is the video-to-world-model bridge: unlike robotics-native peers, it monetizes video generation (Vidu) while converting the underlying physics understanding into manipulation models (Motus), and it open-sources the first-generation world model to seed the ecosystem [5][7].

4. Unique Observations

  • "Video quality = physics understanding" is the core bet. ShengShu's stated rationale for moving from Vidu to Motus is that strong video generation already requires modeling physical laws, so the transition to physical-world action is a continuation rather than a pivot — a thesis shared with minimax's H3 but argued most explicitly here [5].
  • The dexterous-manipulation benchmark sweep. MotuBrain's RoboTwin 2.0 randomized score (96.1%, the only model >95) and Motus2's "self-evolving" value-model loop position ShengShu at the frontier of the manipulation problem most VLA players still avoid [6][7].
  • Local coverage timed with Motus2's release. Jimmy's Sept 2026 research (09-22/09-23) asked whether Minimax H3 could functionally cover Vidu + Motus, concluding H3 was close on coverage but lagged on fine-detail output — an early read on whether one world model can serve both video and robot customers [1][2].

5. Financials / Funding

  • Founded: March 2023, Beijing [5].
  • Series B: ~¥2B (≈460億円) led by Alibaba Cloud, at a ~¥12B valuation (unicorn status) [5].
  • Further round: a new ~$500M (5亿美元) raise reported July 2026 [8].

6. People & Relationships

  • CEO / co-founder: Luo Yihang (骆怡航) [5][7].
  • Academic partner: Tsinghua University — the 朱军 (Zhu Jun) team behind the L1–L5 world-model roadmap — co-developed Motus2 [7].
  • Investors: Alibaba Cloud (Series B lead) [5].
  • Competitors / peers: galaxea, physical-intelligence, minimax, agibot.
Last compiled: 2026-09-28