Unified embodied intelligence

EWAM Emergent Depth‑Wise Specialization
in a Unified Embodied Model

From semantic understanding, through visual foresight, to action formation.

Network depth Β· shallow to deep
01

Shallow layers

Semantic understanding

Action queries retrieve instruction and current-scene context from the VL expert.
02

Intermediate layers

Visual foresight

Attention shifts toward predicted future-frame representations.
03

Deep layers

Action formation

Action self-attention dominates as the model refines motor trajectories.
01 / Overview

One model.
Three forms of computation.

The vision-language expert contributes semantics and the video expert contributes prediction.
EWAM lets one action stream read both at every depth.

Model

One action-centric model

Action tokens read semantic, visual, predicted-future, and action information at every layer.

Finding

An emergent handoff

Without layer-wise supervision, attention moves from VL to predicted future to action across depth.

Results

Strong in sim and on robots

Surpasses VLA, WAM, and hybrid baselines; human data and subtask supervision add further gains.

Where does action attention go?

Existing models let one expert dominate action attention.
EWAM hands attention off from vision-language, to predicted future, to action as depth grows.

Depth-wise handoff
02 / Model & training

Action-centric asymmetric attention

The action stream can read every expert. Perceptual queries remain within their own streams.

VL

Vision-language expert

Qwen3-VL features encode the instruction and current image into layer-aligned semantic representations.

V

Video expert

A Wan2.2 video expert evolves noisy future-video latents into predictive visual context.

A

Action expert

Noisy action chunks query semantic, observed, predicted, and action keys in a shared forward process.

Training recipe

Pretraining regime B 2,084 hours

Human egocentric video

VITRA-1M, EgoDex, and Xperience, with wrist motion expressed as camera-relative end-effector displacements so human wrists and robot end-effectors share one task-space layout.

How human data is used β†’
Task-specific post-training 2 heads subtask text Β· phase index

Long-horizon awareness

A lightweight decoder predicts the active subtask and an MLP head regresses the phase index from the action stream. Both losses are added only during post-training; no external planner is used.

See the long-horizon results β†’
03 / Emergent specialization

A learned handoff across network depth

The architecture does not prescribe shallow, middle, or deep roles. The organization appears during training, replicates across tasks, and remains stable over denoising steps.

Layers 0–975.4%

Shallow understanding

VL tokens receive most action-query attention, grounding the instruction in the current scene.

Layers 10–19Layer 13

Intermediate foresight

Raw future-frame attention overtakes VL-image attention as prediction becomes the dominant visual source.

Layers 20–2959.5%

Deep action formation

Action self-attention dominates while concise semantic and predictive context is retained.

How action attention shifts with depth

Each map shows the share of action-query attention on one source, per layer (rows) and inference step (columns).
VL dominates shallow layers, video peaks in the middle, and action self-attention takes over in deep layers.

Which visual source is read at each depth?

Splitting visual attention into the current image and the predicted future frames.
Predicted frames replace the current image as the main visual source from about layer 13, on one task (a) and across all 50 RoboTwin tasks (b).

Learned during training

The handoff forms within the first few hundred steps

Future-frame minus VL-image attention (pp)

Shallow Middle Deep -60 -40 -20 0 20 40 0 5 10 15 20 25 29 Transformer layer ↑ future > image ↓ image > future 50 steps 150 steps 1,000 steps 40,000 steps
  1. 50 steps

    VL-image leads at almost every layer; no division yet.

  2. 150 steps

    Future frames lead from layer 8 onward.

  3. 40,000 steps

    Shallow layers return to the image; future leads from layer 13 onward.

One training run without robot pretraining. The attention mask sets which sources are reachable, not the depth at which each is used. Hover a step to highlight it.

Action generation depends on it

Masking keys at inference, by source and depth

Success (%)
Full model
92.80
All video Β· L0–9
91.98
VL Β· L0–9
76.53
Condition Β· L10–19
81.40
Future Β· L10–19
3.55
All video Β· L10–19
0.25
All video Β· L20–29
65.82
  1. Shallow Β· L0–9VL matters

    βˆ’16.27 without VL; only βˆ’0.82 without video.

  2. Middle Β· L10–19Future is critical

    βˆ’89.25 when predicted-future keys are masked.

One fixed robot-pretrained checkpoint on RoboTwin Randomized (50 tasks Γ— 100 episodes), masked at inference without retraining.

Cross-task replication 50 / 50

Every RoboTwin task exhibits shallow VL dominance.

Denoising stability 0.998

Mean pairwise correlation of attention profiles across measured denoising steps.

04 / Results

Strong policies across simulation and robots

EWAM is evaluated under complementary protocols: RoboTwin 2.0 clean-to-random and in-domain adaptation, the four LIBERO suites, and seven tasks on three real-world embodiments. Click a card to open the full table.

Real-world evaluation

Seven tasks on three embodiments

Ο€0.5EWAM
FrankaStack Bowls
70
80
FrankaObject into Box
80
90
DobotPour Water
30
70
DobotTidy Desk
20
80
DobotTowel
40
40
G1-DKettle Pouring
60
70
G1-DPour Beans
60
80

Success rate (%) on physical robots. EWAM is higher than Ο€0.5 on six tasks and ties on the Dobot towel task.

Long-horizon progress

Subtask-phase supervision

w/o subtaskw/ subtask
Instructed Ranking
83.5
91.0
Instructed Stacking
44.5
58.5
Ranking & Stacking
1.5
16.0

Full-episode success (%) over 200 episodes per instruction-conditioned block task.

05 / Human egocentric data

Low-cost human data for transfer and robustness

EWAM uses human data in two separate ways: large-scale public egocentric video for pretraining, and self-collected task demonstrations retargeted to the robot for co-training.

Route 1

Human pretraining improves cross-embodiment transfer

About 2,084 hours of VITRA-1M, EgoDex, and Xperience video are reduced to one camera-relative wrist-motion action space. Policies trained on four RoboTwin robots are then evaluated on the held-out Aloha-Agilex-2; the compared policies differ only in whether they start from human pretraining.

Held-out robot embodiment

Human pretraining improves transfer

Without human pretrainingWith human pretraining
60
40
20
33.5
44.3
40K
37.2
53.4
50K
40.1
58.9
80K
36.2
66.9
120K

Aggregate success over 50 RoboTwin 2.0 tasks on the unseen Aloha-Agilex-2 embodiment.

2,084 h

Public egocentric video, with wrist motion as the action target.

40.1 β†’ 66.9%

Best held-out-embodiment success without and with human pretraining.

36.2 β†’ 66.9%

Matched budget of 120K robot-training steps: a 30.7-point gain.

40 β†’ 60%

G1-D pouring under the same 60 robot + 30 human co-training mixture, robot- vs. human-pretrained.

Route 2

Egocentric co-training improves real-robot robustness

For co-training, we record task-specific demonstrations with a headset and wrist trackers, convert them into robot-compatible observation–action trajectories, verify them by replay, and mix them with robot episodes.

1Capture
A headset records fisheye RGB, depth, and camera motion, while two wrist trackers measure metric 6-DoF wrist motion.
2Convert
Hand reconstruction, tool alignment, and grasp-state estimation turn each human recording into robot trajectories paired with the captured images.
3Verify
The same pour as a human recording, a simulation replay, and a real G1-D replay. Robot rows replay the converted trajectories, not policy rollouts.
Replay walkthrough: one human demonstration and its simulation and G1-D replays, aligned phase by phase from grasp to release.
4Co-train
Real-robot robustness

Egocentric data scaling under tabletop and cup changes

Robot only60 robot + egocentric
60
40
20
0
30robot only
10
60robot only
20
90robot only
20
+30egocentric
30
+60egocentric
40
+120egocentric
60
+1,000egocentric

All configurations omit EWAM human and robot pretraining and are evaluated over 10 physical trials.

30 episodes

Adding 30 egocentric episodes to 60 robot episodes reaches 20%, the same as 90 robot-only episodes, at a much lower collection cost.

10% β†’ 60%

Adding 30 to 1,000 egocentric episodes keeps raising success under tablecloth and cup changes, without saturating.

06 / Demonstrations

Real-world manipulation across embodiments

Rollouts of EWAM across embodiments and tasks.

01TowelDobot dual-arm
02Tidy DeskDobot dual-arm
03Pour WaterDobot dual-arm
04Stack BowlsFranka
05Pour BeansUnitree G1-D
06Pour Beans Β· Tablecloth ChangeUnitree G1-D
07 / Authors

Authors

Core Contributors

Hao Wang1,2, Jiajun Wen3, Jingzhi Liu1, Shuoshuo Xue1, Zhiliang Chen2,4

Contributors

Min Lin1, Yicheng Chang1, Xiaoyu Guo5, Yukang Zhuo1, Zheng Chong1, Yunshuang Nie3, Jian Zhang6, Weijia Liufu1, Qingman Wu1, Heming Xu7, Bingchang Song8, Dantong Wu9, Zhiyuan Wang9

Guidance Group

Hang Xu9, Jianhua Han9, Bokui Chen3, Shen Zhao1

Project Leaders

Rui Li9, Xiaodan Liang1,9

  1. Sun Yat-sen University
  2. Beijing Zhongguancun Academy
  3. Tsinghua University
  4. Fudan University
  5. South China University of Technology
  6. Mohamed bin Zayed University of Artificial Intelligence
  7. Northwest University, Xi'an
  8. Southern University of Science and Technology
  9. Yinwang Intelligent Technology Co. Ltd.
08 / Cite

Read and cite EWAM

Read the full technical report or copy the citation.

Technical report
@article{wang2026ewam,
  title   = {{EWAM}: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action},
  author  = {Wang, Hao and Wen, Jiajun and Liu, Jingzhi and Xue, Shuoshuo and Chen, Zhiliang and Lin, Min and Chang, Yicheng and Guo, Xiaoyu and Zhuo, Yukang and Chong, Zheng and Nie, Yunshuang and Zhang, Jian and Liufu, Weijia and Wu, Qingman and Xu, Heming and Song, Bingchang and Wu, Dantong and Wang, Zhiyuan and Xu, Hang and Han, Jianhua and Chen, Bokui and Zhao, Shen and Li, Rui and Liang, Xiaodan},
  journal = {arXiv preprint arXiv:2609.39973},
  year    = {2026}
}

β–Ά

Video will be added here.

BibTeX copied