StairVLA

StairVLAStage-Aware Hierarchical Action Generation for Vision-Language-Action Models

Partially denoise a long-horizon plan, reuse it, and finish each action chunk from the latest observation.

Shangyuan Yuan1 Xinda Qi1,2 Yujiang Pu1 Wenliang Guo1 Xiaobo Tan1

1Michigan State University    2Ant Technology U.S., Inc.

A high-level VLA runs the early denoising steps once and caches a long-horizon, partially denoised trajectory. A lightweight refiner finishes each local action chunk from the latest observation, so the expensive VLA call is shared across several control cycles.
+1.3 ptsLIBERO success, 96.5% → 97.8% (GR00T-base)
2.6× fasterlatency per action chunk, 115.0 → 44.2 ms
81.2% · 60.0%real-robot success on Fruit25 and PushBlock, best of three methods

Abstract

Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions. These action heads typically process the denoising trajectory in a largely uniform manner. However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment. Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement. A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations. This design amortizes expensive high-level VLA computation while preserving frequent closed-loop correction. On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk. More broadly, across two VLA backbones, simulation benchmarks, and real-robot tasks, StairVLA consistently reduces inference cost while maintaining strong task performance.

Motivation

The conditioning focus of a generative action head changes as denoising proceeds: instruction and image jointly shape the coarse trajectory early on, and the current image matters more near the end. Task competence is largely in place before denoising completes.

Stage-dependent modality reliance, attention analysis across denoising progress, and task success versus denoising progress
(A) Conditioning focus shifts across denoising stages. (B) Relative attention to image tokens grows toward the end of denoising. (C) Success rate when executing partially denoised actions rises well before denoising completes.

Method

StairVLA uses the partially denoised trajectory as the interface between two levels.

StairVLA framework: high-level generation, low-level refinement, and action execution over time
The high-level VLA turns noise into a partially denoised long-horizon trajectory at t0. At each later step, the refiner takes the matching segment and the newest observation and produces an executable action chunk.
1

Early denoising at the high level

A VLA (Qwen3-VL with a flow-matching head) denoises a long horizon (H = 32) only up to an intermediate progress α, keeping the task intent encoded by instruction and image.

2

Reuse across control cycles

The partially denoised trajectory is cached and reused for M refinement cycles, so the expensive VLA call is amortized over many executed actions.

3

Late refinement from fresh observations

A lightweight DiT refiner completes the remaining denoising for each chunk, conditioned on a SigLIP encoding of the latest image, compressed VLA features, and neighboring trajectory context.

Refiner architecture: SigLIP visual encoder, token compression, partially denoised actions with left and right context, and a cross-attention DiT
The refiner. Current-observation tokens, compressed high-level VLM tokens, and the query chunk with its left and right context feed a DiT that alternates cross- and self-attention and predicts the remaining flow-matching velocity.

Results

LIBERO

Success rate (%) over 500 trials per suite with one model trained on all four suites. Latency is per executable action chunk on one NVIDIA A100, including the amortized high-level cost.

LIBERO success rate versus latency StairVLA moves both backbones up and to the left: GR00T-base from 96.5% at 115.0 ms to 97.8% at 44.2 ms; pi-base from 97.8% at 183.3 ms to 98.3% at 44.5 ms. 96 97 98 0 50 100 150 200 Latency per action chunk (ms) · lower is better LIBERO avg. success (%) · higher is better 2.6× faster · +1.3 pts 4.1× faster · +0.5 pts StarVLA-GR00T StairVLA (GR00T-base) StarVLA-π StairVLA (π-base)
For both backbones, StairVLA is more accurate and several times faster than the StarVLA model it builds on. The table below lists the same numbers.
MethodSpatialObjectGoalLongAvg.Latency (ms)
StarVLA-GR00T†97.898.897.492.096.5115.0
StairVLA (GR00T-base)98.099.697.895.897.844.2
StarVLA-π†99.299.097.295.897.8183.3
StairVLA (π-base)99.699.096.897.498.344.5

† Success rates reported by StarVLA. Comparisons with other hierarchical and coarse-to-fine VLAs are in the paper.

Ablations

SettingAvg.Δ
StarVLA-GR00T96.5–
High-level policy only (H = 32)92.4−4.1
StairVLA w/o action context97.2+0.7
StairVLA97.8+1.3

Running the long-horizon policy alone loses accuracy; the observation-conditioned refiner recovers it, and the temporal action context adds a further gain.

Latency and success rate versus reuse horizon M for Libra-VLA, top-only, and StairVLA
Reusing each high-level trajectory for more refinement cycles (larger M) lowers latency while success stays high.

LIBERO-Plus

Robustness to seven perturbation types, zero-shot from LIBERO and after fine-tuning on LIBERO-Plus.

SettingCameraRobotLanguageLightBackgroundNoiseLayoutAvg.
Zero-shot43.754.582.593.490.567.874.070.4
Fine-tuned95.548.483.796.894.995.477.483.7

Real-Robot Experiments

An AgileX PiperX arm with a fixed top camera and a wrist camera. Fruit25 has 24 pick-and-place tasks plus a multi-object sorting task (8 Hz); PushBlock is a contact-rich pushing task (20 Hz). All methods use Qwen3-VL-2B and run on the same RTX A4000.

Real-world success rate and latency of StarVLA-GR00T, StarVLA-pi, and StairVLA on Fruit25 and PushBlock
Success rate and model latency on the real robot.

PushBlock

Top camera, 4× speed. Red border model inference, green border action execution; the counter shows executed action chunks.

Harder start, similar layout: StairVLA pushes the block into the target; StarVLA-GR00T does not.
Easier start: both succeed, StairVLA at about 36 s and StarVLA-GR00T near the 50 s limit.

Fruit25

StairVLA on three language-conditioned tasks (2× speed).

Full video

Supplementary video: PushBlock comparisons under two initial configurations and Fruit25 examples.

BibTeX

@misc{yuan2026stairvla,
  title         = {StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models},
  author        = {Yuan, Shangyuan and Qi, Xinda and Pu, Yujiang and Guo, Wenliang and Tan, Xiaobo},
  year          = {2026},
  eprint        = {2610.07756},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2610.07756}
}

StairVLA is built on StarVLA.