← VLA Atlas

Motus RoboTwin 87% · +45% vs π0.5 · +15% vs X-VLA

Tsinghua · Mixture-of-Transformer + UniDiffuser scheduler · 一个模型同时是 WM / VLA / IDM / video-gen / joint prediction 五种

MotusarXiv 2512.13030)解决的问题:现在的具身系统被拆成 understanding / world model / control 三家孤立训练,不能共享参数、不能吃海量异构数据。Motus 用 MoT(Mixture-of-Transformer)把三个专家(understanding、video generation、action)塞进同一个模型;用 UniDiffuser-style scheduler 通过控制"哪些 token 已知 / 待生成",在一个模型内切换五种运行模式

关键 trick 有两个:(1) 用光流作为 latent action,规避动作标注稀缺问题;(2) 三阶段训练 + 六层数据金字塔,让不同标注密度的数据都能贡献。

结果在 RoboTwin 2.0 大规模 randomized 多任务上打过所有 baseline —— Motus 87.02% vs π₀.₅ 43.84%(差 +45%),且真机 AC-One / Agilex-Aloha-2 同样明显胜出。

五种运行模式(一个模型五种用法)

  1. World Model:给动作 + 当前帧 → 预测未来帧
  2. VLA:给指令 + 当前帧 → 预测动作
  3. IDM(Inverse Dynamics Model):给相邻两帧 → 反推动作
  4. Video Generation:给指令 → 生成视频
  5. Video-Action Joint Prediction:同时输出未来帧 + 动作序列

UniDiffuser 提供的核心机制是"任意子集条件"—— 通过 masking,模型能学到任何一种 conditional 分布,切换五种任务只是改 mask,不换权重。

RoboTwin 2.0 主结果(论文 Table 2)

50 个 RoboTwin 任务的多任务联合训练:2500 clean demos(50/task) + 25000 randomized demos(500/task),所有 baseline 从各自 pretrained checkpoint fine-tune 40k steps,测试 100 trials/task。摘录 20 个代表性任务:

Taskπ₀.₅ (Rand.)X-VLA (Rand.)w/o Pre (Rand.)Stage1 (Rand.)Motus (Rand.)
Place Dual Shoes7%88%80%94%87%
Move Stapler Pad18%73%37%68%85%
Stack Blocks Two56%87%94%99%98%
Scan Object38%36%50%69%66%
Pick Dual Bottles6%36%68%17%90%
Turn Switch6%61%60%64%78%
Pick Diverse Bottles3%36%62%18%91%
Hanging Mug3%27%10%25%38%
Put Bottles Dustbin9%77%33%24%79%
Open Microwave37%71%82%84%91%
50 任务平均(Randomized)43.84%72.84%77.00%81.86%87.02%
50 任务平均(Clean)42.98%72.80%72.8%82.86%88.66%

真机实验 · AC-One 平台(论文 Table 3)

Taskπ₀.₅w/o PreMotus
Fold Towel4114.5
Brew Coffee using Coffee Maker0062
Get Water from Dispenser30836
Place Cube into Plate4660100
Place Cube into Plate (OOD)281975
Grind Coffee Beans8092
Pour Water from Kettle to Flowers5565
Touch Instructed Keyboard010082.5
Put Bread into Oven124042
Average14.7925.8663.22

评估用部分成功率(partial success rate),每个任务分解为子目标,达成子目标得部分分。这对长时任务更公平。

Agilex-Aloha-2 平台

Taskπ₀.₅w/o PreMotus
Fold Towel27.5039
Get Water from Dispenser62896
Pour Water from Kettle to Flowers454047.5
Touch Instructed Keyboard72.58580
Put Bread into Oven36034
Average48.6026.6059.30

训练配方

差异化定位

Motus 是"World Model + VLA 统一"路线里目前数字最有说服力的。它和其他统一方案的对比:

相关

Genie Envisioner
另一条"世界模型 + VLA 一体化"路线
UniVLA
同为统一路线的 pure-AR 变体
X-VLA(被超越对象)
Motus 相对提升 +15%
π₀.₅(被超越对象)
Motus 相对提升 +45%

参考链接

arXiv 2512.13030
Motus: A Unified Latent Action World Model · 2025-12-15