Motus RoboTwin 87% · +45% vs π0.5 · +15% vs X-VLA
Tsinghua · Mixture-of-Transformer + UniDiffuser scheduler · 一个模型同时是 WM / VLA / IDM / video-gen / joint prediction 五种
Motus(arXiv 2512.13030)解决的问题:现在的具身系统被拆成 understanding / world model / control 三家孤立训练,不能共享参数、不能吃海量异构数据。Motus 用 MoT(Mixture-of-Transformer)把三个专家(understanding、video generation、action)塞进同一个模型;用 UniDiffuser-style scheduler 通过控制"哪些 token 已知 / 待生成",在一个模型内切换五种运行模式。
关键 trick 有两个:(1) 用光流作为 latent action,规避动作标注稀缺问题;(2) 三阶段训练 + 六层数据金字塔,让不同标注密度的数据都能贡献。
结果在 RoboTwin 2.0 大规模 randomized 多任务上打过所有 baseline —— Motus 87.02% vs π₀.₅ 43.84%(差 +45%),且真机 AC-One / Agilex-Aloha-2 同样明显胜出。
五种运行模式(一个模型五种用法)
- World Model:给动作 + 当前帧 → 预测未来帧
- VLA:给指令 + 当前帧 → 预测动作
- IDM(Inverse Dynamics Model):给相邻两帧 → 反推动作
- Video Generation:给指令 → 生成视频
- Video-Action Joint Prediction:同时输出未来帧 + 动作序列
UniDiffuser 提供的核心机制是"任意子集条件"—— 通过 masking,模型能学到任何一种 conditional 分布,切换五种任务只是改 mask,不换权重。
RoboTwin 2.0 主结果(论文 Table 2)
50 个 RoboTwin 任务的多任务联合训练:2500 clean demos(50/task) + 25000 randomized demos(500/task),所有 baseline 从各自 pretrained checkpoint fine-tune 40k steps,测试 100 trials/task。摘录 20 个代表性任务:
| Task | π₀.₅ (Rand.) | X-VLA (Rand.) | w/o Pre (Rand.) | Stage1 (Rand.) | Motus (Rand.) |
|---|---|---|---|---|---|
| Place Dual Shoes | 7% | 88% | 80% | 94% | 87% |
| Move Stapler Pad | 18% | 73% | 37% | 68% | 85% |
| Stack Blocks Two | 56% | 87% | 94% | 99% | 98% |
| Scan Object | 38% | 36% | 50% | 69% | 66% |
| Pick Dual Bottles | 6% | 36% | 68% | 17% | 90% |
| Turn Switch | 6% | 61% | 60% | 64% | 78% |
| Pick Diverse Bottles | 3% | 36% | 62% | 18% | 91% |
| Hanging Mug | 3% | 27% | 10% | 25% | 38% |
| Put Bottles Dustbin | 9% | 77% | 33% | 24% | 79% |
| Open Microwave | 37% | 71% | 82% | 84% | 91% |
| 50 任务平均(Randomized) | 43.84% | 72.84% | 77.00% | 81.86% | 87.02% |
| 50 任务平均(Clean) | 42.98% | 72.80% | 72.8% | 82.86% | 88.66% |
真机实验 · AC-One 平台(论文 Table 3)
| Task | π₀.₅ | w/o Pre | Motus |
|---|---|---|---|
| Fold Towel | 4 | 1 | 14.5 |
| Brew Coffee using Coffee Maker | 0 | 0 | 62 |
| Get Water from Dispenser | 30 | 8 | 36 |
| Place Cube into Plate | 46 | 60 | 100 |
| Place Cube into Plate (OOD) | 28 | 19 | 75 |
| Grind Coffee Beans | 8 | 0 | 92 |
| Pour Water from Kettle to Flowers | 5 | 5 | 65 |
| Touch Instructed Keyboard | 0 | 100 | 82.5 |
| Put Bread into Oven | 12 | 40 | 42 |
| Average | 14.79 | 25.86 | 63.22 |
评估用部分成功率(partial success rate),每个任务分解为子目标,达成子目标得部分分。这对长时任务更公平。
Agilex-Aloha-2 平台
| Task | π₀.₅ | w/o Pre | Motus |
|---|---|---|---|
| Fold Towel | 27.5 | 0 | 39 |
| Get Water from Dispenser | 62 | 8 | 96 |
| Pour Water from Kettle to Flowers | 45 | 40 | 47.5 |
| Touch Instructed Keyboard | 72.5 | 85 | 80 |
| Put Bread into Oven | 36 | 0 | 34 |
| Average | 48.60 | 26.60 | 59.30 |
训练配方
- MoT 架构:Mixture-of-Transformer(不是 Mixture-of-Experts)· 三路专家:understanding / video-gen / action
- Latent action 用光流:像素级 delta action,来源是通用视频,规避了动作标注稀缺
- 三阶段训练:Stage 1 单模式训 → Stage 2 联合训 → Stage 3 微调
- 六层数据金字塔:从纯视频、视频+光流、视频+动作 逐级加标 —— 让不同标注密度的数据都能用
- Ablation(Fig 6):w/o Pretrain 77% · Stage1-only 82% · Full Motus 87%(RoboTwin 平均) —— 两阶段都必要
差异化定位
Motus 是"World Model + VLA 统一"路线里目前数字最有说服力的。它和其他统一方案的对比:
- vs UniVLA:UniVLA 是纯自回归 token 统一,Motus 是 MoT + Diffusion;Motus 更能利用大规模无动作标注视频(因为 latent action = optical flow)
- vs Genie Envisioner:GE 是 GE-Base + GE-Act 两个独立模型,Motus 是一个模型五种角色;GE 更工程化(带 GE-Sim),Motus 更架构统一
- vs NORA-1.5:NORA 是 flow-matching action expert + DPO 后训练,Motus 强调架构层面(MoT)而非训练目标