FUTURE-AWARE ROBOT LEARNING

SLIP–VLA

Single-Step Latent Imagination for Policy Learning
in Vision-Language-Action Models

Tianfu Li1,*Haoxuan Xu2,*Wenbo Chen1,*Haitian Li3Changchuan Yang4Xinhu Zheng1Jun Ma1Yuan Liu2Lujia Wang1Haoang Li1
1 The Hong Kong University of Science and Technology (Guangzhou)2 The Hong Kong University of Science and Technology3 Nanyang Technological University4 Zhejiang University

* Equal contribution

One forward pass. A future ready for action.
12ms

Latent imagination

Single World DiT pass
98.9%

LIBERO

Average success rate
83.2%

LIBERO-Plus

Under distribution shifts
91.8%

RoboTwin 2.0

Tied best average in the paper

THE IDEA

Imagine the future.
Act in a single step.

What if a robot could anticipate how a scene will evolve without waiting for an entire video to be generated?

SLIP-VLA produces temporally dense future latents with a single denoising update. We shape these representations to preserve future geometry, task-relevant semantics, and the effects of robot actions, then inject them directly into the action policy.

The auxiliary teachers and dynamics models are used only during training. At inference, future imagination takes 12 ms; overall inference latency is 181 ms.

Read the abstract

Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive iterative denoising, while one-step alternatives can underperform their multi-step counterparts. To reconcile efficient future modeling with strong action performance, we present SLIP-VLA, a policy learning framework that equips VLA models with a Single-Step Latent Imagination for future-aware action prediction. SLIP-VLA obtains temporally dense future latent representations with a single denoising update, and we improve the perceptual sufficiency of these representations by aligning intermediate latents with future geometric and semantic features. We further improve their control sufficiency through action-conditioned latent world modeling and inverse dynamics modeling, explicitly coupling latent transitions with robot actions. Our SLIP-VLA achieves state-of-the-art performance across diverse simulation benchmarks and real-world manipulation tasks, while its single-step latent imagination takes only 12 ms.

01 / REAL-WORLD DEMONSTRATIONS

From imagination
to physical interaction.

Six task conditions. Four camera views each.
Explore closed-loop manipulation on the
AgileX COBOT Magic platform.

Success rates are from Table VIII of the paper, with 50 trials per condition. The videos show selected successful rollouts.

02 / HOW IT WORKS

A useful future.
Without the wait.

Four future slots. One World DiT pass.
Future representations guide the action
decoder at multiple layers.

01Observe

Images + language

02 · SINGLE STEPImagine

4 future latent slots

12 ms
03Act

16-step action chunk

SLIP-VLA architecture: the VLM conditions a single-step World DiT, perceptual and control objectives shape future latents during training, and layer-wise injection conditions the Action DiT.
SLIP-VLA architecture. Geometric and semantic teachers, forward dynamics, and inverse dynamics provide training-time supervision only.
A

Perceptual sufficiency

Future VGGT-Ω geometry and DINOv3 semantics shape intermediate representations to preserve where interactions happen and which objects matter.

B

Control sufficiency

Forward dynamics predicts latent evolution from actions. Inverse dynamics recovers actions from latent transitions, making the imagined future informative for control.

03 / SIMULATION RESULTS

Stronger policies.
Across harder settings.

Selected rollouts alongside StarVLA.
Explore standard manipulation, visual
shifts, and bimanual coordination.

Video comparisons use StarVLA as the baseline. Quantitative results below follow the manuscript, which evaluates a separate set of baselines. Selected rollouts are illustrative, not aggregate evaluations.

ROBUSTNESS / LIBERO-PLUS

83.2% average success

+3.9 points over the strongest compared method in Table II.

SLIP-VLA83.2
LaMP79.3
OpenVLA-OFT69.6
Fast-WAM-IDM68.8
Fast-WAM-Joint68.7

Selected methods from Table II. Values are success rates (%).

ABLATION / LATENT SHAPING

Both objectives matter.

Perceptual and control shaping work best together.

Ablation results from Table IV, success rates in percent
VariantLIBEROLIBERO-Plus
Naive single-step97.779.0
+ Perceptual98.381.4
+ Control98.582.3
+ Both (SLIP-VLA)98.983.2

Table IV. Same single-step setting across variants.

Qualitative evidence that shaped single-step latents preserve future geometry, task-relevant semantics and action-dependent dynamics.
What the latent learns. More coherent future geometry, concentrated semantic responses, and sensitivity to action changes (Figure 3).
Denoising step study: one step achieves 83.2% LIBERO-Plus success; two, three and four steps achieve 83.1%, 82.9%, and 82.9%.
One step is enough. Additional denoising does not improve policy success in the reported experiment (Figure 4).

04 / CITATION

Build on SLIP-VLA.

Published on arXiv.
Use the citation below
when referring to our work.

@misc{li2026slipvlasinglesteplatentimagination,
  title = {{SLIP-VLA}: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models},
  author = {Tianfu Li and Haoxuan Xu and Wenbo Chen and Haitian Li and Changchuan Yang and Xinhu Zheng and Jun Ma and Yuan Liu and Lujia Wang and Haoang Li},
  year = {2026},
  eprint = {2609.33575},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.33575},
}

Use each player's controls to inspect a view. Press Esc to close.