A unified diagnostic benchmark

VLA-GAUGE:A Systematic Evaluation Framework for Generalization and Robustness of VLA Models

The first unified diagnostic framework for systematically evaluating zero-shot generalization and perturbation robustness in VLA models

Overview

Benchmark comparison

Evaluation that closes the diagnostic loop

Existing benchmarks lack: ① fine-grained failure diagnosis across vision-language understanding, action representation, and control execution; ② unified capability-level In-distribution (ID)/Out-of-distribution (OOD) analysis; ③ realistic combinatorial testing; and ④ support for model refinement and re-evaluation. VLA-GAUGE is the first systematic framework to jointly evaluate and refine VLA policies, including capability-oriented paired ID/OOD evaluation, progressive combinatorial perturbation evaluation and diagnosis-guided models refinemen

6,300Task Instances
3,089Objects
Full support Full SupportPartial support Partial SupportNo support No Support
Comparison of VLA-GAUGE with existing robot benchmarks
BenchmarkFocusTasksObjectsFine-grainedID/OODRobustnessOptimization
LIBEROKnowledge Transfer13075Partial supportNo supportNo supportNo support
Behavior-1KLong-horizon + Dataset1,0009,331Partial supportNo supportNo supportNo support
COLOSSEUMRobustness20147Partial supportPartial supportPartial supportPartial support
SimplerEnvGeneralization817Partial supportPartial supportPartial supportNo support
VLA-TestRobustness474Partial supportFull supportPartial supportNo support
VLABenchGeneralization + Long-horizon1002,164Partial supportPartial supportPartial supportNo support
VLA-ArenaRobustness170122Partial supportPartial supportPartial supportNo support
AGNOSTOSGeneralization23112No supportFull supportNo supportNo support
LIBERO-XRobustness60096Partial supportPartial supportPartial supportNo support
LIBERO-PRORobustness4077Partial supportFull supportPartial supportNo support
LIBERO-PlusRobustness10,030911Partial supportFull supportPartial supportPartial support
VLA-GAUGEGeneralization + Robustness6,3003,089Full supportFull supportFull supportFull support

Benchmark construction

Four modules from raw meshes to evaluation-ready data

Assets, tasks, perturbations and trajectories all come out of one toolchain. Each module below lists its stages, an example of what it emits, and the constraints that matter in practice.

01Object Asset Generation

From external meshes to simulation-ready objects

The input is a scanned or downloaded mesh; the output is a registered simulation asset carrying collision geometry and anchor sites.

Input MeshGeometry and texturesNormalizeOne folder per objectCollision MeshConvex decompositionRegistrationAnchor sites, then the inventory
Four stages from a raw mesh to a simulation object
Raw mesh, textured appearance, convex collision hulls, and the simulation object with anchor sites
  • Input: one directory per object, holding the geometry, material and texture files.
  • Collision: convex decomposition is automatic; containers also receive a containment region.
  • Registration: each asset becomes an object class that scenes and tasks reference by name.
  • Manual step: scale and origin are often off, so scale and anchor poses need calibration.

02Scene and Task Construction

From the object inventory to executable task definitions

Available objects and scenes are confirmed first, then the scene layout and the task definition are generated, and initial states are sampled for every task.

InventoryObjects and scenes on handScene LayoutObjects and candidate regionsTask DefinitionTarget object and goal stateInitial StatesSampled states per task
Tabletop scene with candidate regions and a marked target object
The scene layout defines candidate regions on the table (green); the task definition marks the target object (white box)
  • Scene layout: fixtures and objects are drawn at random, with non-overlapping candidate regions.
  • Task definition: an LLM reads the scene state and emits the target object and goal state.
  • Initial states: several states are sampled per task, shared by collection and evaluation.
  • Manual review: instruction wording, goal states and object poses all need a pass.

03Perturbation Generation

Three families of controllable perturbations, stackable into compositions

Perturbations are grouped by when they take effect: applied at runtime, baked into the scene, or baked into the object asset.

Runtime-levelCamera, init state, noise, image, actionScene-levelLight direction, diffuse, specular, shadowObject-levelSize, orientation, color, distractorsCompositionSeveral factors at once
One scene rendered under six different perturbations
Six variants of one scene: baseline, lighting, camera viewpoint, distractors, target appearance, and sensor noise
  • Runtime: camera viewpoint, initial state, sensor noise, image transforms, action noise and dynamic lighting.
  • Scene: light direction, diffuse, specular and shadow are rewritten into reproducible variants.
  • Object: size, orientation and color are resampled, and distractors are added or removed.
  • Combinatorics: counts grow fast once factors stack, so cap the number of combinations.

04Trajectory Replay

Teleoperated collection, batch replay and dataset export

Demonstrations are collected by teleoperation, replayed so observations can be re-rendered and validated, then filtered and packaged into a dataset.

TeleoperationImages, actions, physics stateReplay CheckRe-run actions, re-render viewsFrame FilterDrop idle frames and failuresPackagingReady for training
Keyframes, end-effector trajectory and wrist camera view of one demonstration
Top: keyframes of one demonstration. Bottom: the end-effector trajectory and the wrist camera view
  • Collection: keyboard or space-mouse teleoperation, recording images, actions and physics state.
  • Validation: success must hold ten frames, ten idle frames are appended, final reward above 0.99.
  • Replay: actions are re-executed to re-render color, depth and end-effector states.
  • Export: videos for visual inspection; with idle frames removed the data is training-ready.

Leaderboards

Three evaluation views of model reliability

Compare competitive LIBERO VLA models across zero-shot generalization, single-factor sensitivity, and combinatorial perturbation robustness.

1 / 3

Zero-Shot Generalization Leaderboard

Zero-Shot Score (ZS) for every model across the six capability axes: the mean performance over the ID and OOD tasks within each suite. Switch views to split each score into its in-domain and out-of-domain components, or to read the gap between them directly. Models are divided into small/large scale groups using a 5B-parameter cutoff.

12 entries

Values are overall ZS scores across six capability axes (higher is better). Rows are ranked by average ZS in descending order.

Zero-shot generalization of 12 VLA models across six capability axes, shown as Overall.
Rank
1π0.53BM — Mixed action space60.057.861.847.672.741.896.957.0
2DeepThinkVLA2.9BD — Discrete action space39.139.637.821.833.229.897.033.6
3Evo-10.77BC — Continuous action space30.228.030.78.028.223.694.824.8
4OpenVLA-OFT7BC — Continuous action space38.210.236.416.922.721.397.124.3
5UniVLA9BD — Discrete action spaceHist. — uses observation or action history29.318.723.610.235.923.695.223.6
6X-VLA0.9BC — Continuous action space17.314.79.822.737.335.698.122.9
7CronusVLA7BC — Continuous action spaceHist. — uses observation or action history24.428.929.38.418.222.797.022.0
8RIPT-VLA7BC — Continuous action space30.26.728.45.315.323.197.518.2
9GR00T1.53BC — Continuous action spaceHist. — uses observation or action history15.116.017.34.914.517.393.914.2
10VLA-Adapter-Pro0.5BC — Continuous action space16.414.77.66.210.519.698.512.5
11SimpleVLA-RL7BD — Discrete action space11.12.28.02.213.611.199.18.0
12Nora3BD — Discrete action space11.19.38.42.26.81.887.96.6
  • 1
    π0.53BM — Mixed action space
    57.0
    Vis.
    60.0
    Sem.
    57.8
    CW
    61.8
    Spa.
    47.6
    Act.
    72.7
    Comp.
    41.8
  • 2
    DeepThinkVLA2.9BD — Discrete action space
    33.6
    Vis.
    39.1
    Sem.
    39.6
    CW
    37.8
    Spa.
    21.8
    Act.
    33.2
    Comp.
    29.8
  • 3
    Evo-10.77BC — Continuous action space
    24.8
    Vis.
    30.2
    Sem.
    28.0
    CW
    30.7
    Spa.
    8.0
    Act.
    28.2
    Comp.
    23.6
  • 4
    OpenVLA-OFT7BC — Continuous action space
    24.3
    Vis.
    38.2
    Sem.
    10.2
    CW
    36.4
    Spa.
    16.9
    Act.
    22.7
    Comp.
    21.3
  • 5
    UniVLA9BD — Discrete action spaceHist. — uses observation or action history
    23.6
    Vis.
    29.3
    Sem.
    18.7
    CW
    23.6
    Spa.
    10.2
    Act.
    35.9
    Comp.
    23.6
  • 6
    X-VLA0.9BC — Continuous action space
    22.9
    Vis.
    17.3
    Sem.
    14.7
    CW
    9.8
    Spa.
    22.7
    Act.
    37.3
    Comp.
    35.6
  • 7
    CronusVLA7BC — Continuous action spaceHist. — uses observation or action history
    22.0
    Vis.
    24.4
    Sem.
    28.9
    CW
    29.3
    Spa.
    8.4
    Act.
    18.2
    Comp.
    22.7
  • 8
    RIPT-VLA7BC — Continuous action space
    18.2
    Vis.
    30.2
    Sem.
    6.7
    CW
    28.4
    Spa.
    5.3
    Act.
    15.3
    Comp.
    23.1
  • 9
    GR00T1.53BC — Continuous action spaceHist. — uses observation or action history
    14.2
    Vis.
    15.1
    Sem.
    16.0
    CW
    17.3
    Spa.
    4.9
    Act.
    14.5
    Comp.
    17.3
  • 10
    VLA-Adapter-Pro0.5BC — Continuous action space
    12.5
    Vis.
    16.4
    Sem.
    14.7
    CW
    7.6
    Spa.
    6.2
    Act.
    10.5
    Comp.
    19.6
  • 11
    SimpleVLA-RL7BD — Discrete action space
    8.0
    Vis.
    11.1
    Sem.
    2.2
    CW
    8.0
    Spa.
    2.2
    Act.
    13.6
    Comp.
    11.1
  • 12
    Nora3BD — Discrete action space
    6.6
    Vis.
    11.1
    Sem.
    9.3
    CW
    8.4
    Spa.
    2.2
    Act.
    6.8
    Comp.
    1.8
00.0Best00.0Secondcompared within each row

ID and OOD scenarios are not evenly distributed across capabilities, so the overall score is a weighted combination of the two columns rather than their arithmetic mean.

Join the benchmark

Submit a model for official evaluation

Share the public Hugging Face repository containing your VLA model weights and loading instructions. The evaluation team downloads the model, runs VLA-GAUGE, and updates the verified leaderboard.

  • A readable repository. Weights must sit in a public repository. Gated repositories are accepted, but access has to be approved before the run can start.
  • Instructions we can follow. Say which checkpoint to load and list any dependency the standard setup does not cover.
  • Details for the table. Parameter count, action space, and whether the policy conditions on history — these place the model in the leaderboard groups.
  • A contact address. The team writes to you if the model fails to load or a result needs confirming.

No weights are uploaded through this website.

Submissions are emailed to the evaluation team after the repository check.

Experimental findings

From aggregate score to failure mechanism

Each analysis presents its research question and key finding alongside the corresponding figure and representative rollout videos.

RQ1 (Overall Zero-Shot Generalization Performance): What is the overall zero-shot generalization performance of existing VLA models?

Key evidence supporting the findings

  • 22.3%Mean overall ZS
  • 6.6–57.0%ZS range across 12 models
  • 32.9% → 7.2%Mean ID-to-OOD ZS
  • 71/72Model-capability degraded
  • 32OOD entries scoring 0%

RQ2 (Zero-Shot Performance Distribution and Capability Bottlenecks): What zero-shot performance distribution and capability bottlenecks do existing VLA models exhibit across VLA-GAUGE's six capability dimensions?

Key evidence supporting the findings

  • 13.0%The Lowest overall ZS in spatial reasoning
  • 8/12The lowest or joint-lowest score in spatial reasoning
  • 16.0%The Lowest ID ZS in spatial reasoning
  • 2.8%The Lowest OOD ZS in compositional capability
  • 9/12the OOD scoring 0% in compositional capability

Radar distributions of VLA models across six capability dimensions

4 charts

(a)-(b) Overall ZS and (c)-(d) ID and OOD ZS, both grouped by model scale. Spatial consistently dips inward, while Composite OOD collapses toward the center for most models. Across the evaluated models, parameter scale does not reliably predict zero-shot generalization.

Radar chart of small-scale VLA models across six capability dimensions.
(a) Small Models: Overall ZS
Radar chart of large-scale VLA models across six capability dimensions.
(b) Large Models: Overall ZS
Radar chart of ID and OOD performance for small-scale VLA models.
(c) Small Models: ID vs. OOD
Radar chart of ID and OOD performance for large-scale VLA models.
(d) Large Models: ID vs. OOD

π0.5 · strongest rollout on each axis

6 axes

Each capability axis presents a 2×2 grid of selected successful rollouts, with each clip captioned by its task instruction. Due to layout constraints, some captions under Semantic Grounding, Common Sense and World Knowledge, and Compositional Capability show standard task instructions; the exact scenario-specific instructions are provided in the corresponding files.

Visual Perception

Pick up the mahjong of 5 pin

Put the wine in the basket

Put the wine in the basket

Put the cookies on the plate

Semantic Grounding

Insert a sunflower in a basket

Put the teapot in the basket

Pick up the mahjong of 9 pin

Pick up the eggplant in basket

Common Sense & World Knowledge

Pick up the queen of hearts

Put the cup in the basket

Pick up the mahjong of 4 sou

Put the car in the basket

Spatial Reasoning

Place the farthest fruit on the plate

Pick the first fruit from left to right and put it in the basket

Pick up wine and place it in the left compartment of caddy

put the mug farthest to the chocolate pudding on plate

Action Reasoning

Open the microwave

Pick the small cube and place it in the basket

Turn off the stove

Close the microwave

Compositional Capability

Put the both animal toys in the basket

Turn on the stove and put the black bowl on it

Put both the butter and the cucumber in the basket

I want to put the yellow book away. Could you place it on top of the microwave and close it?

Perturbation Combinations

From Category to Factor Combination

Each PCC and PFC configuration comprises a deliberately selected combination of perturbation factors and their underlying components. Colors denote each factor's high-level perturbation category.

1 / 2

Perturbation Category Combinations

ControlSpatialVisual

PCC selects representative factors from one, two, or all three perturbation categories.

PCC-1.1
Camera Vertical
PCC-1.2
Clutter Distractor Swap
PCC-1.3
Initstate
PCC-2.1
Camera VerticalClutter Distractor Swap
PCC-2.2
Camera VerticalInitstate
PCC-2.3
Clutter Distractor SwapInitstate
PCC-3.1
Camera VerticalClutter Distractor SwapInitstate

Citation

Cite VLA-GAUGE

If you find VLA-GAUGE useful in your research, please consider citing our work.

Reference

BibTeX

Citation metadata will be published with the VLA-GAUGE paper.
Submit your VLA model to the VLA-GAUGE leaderboard|View Submission Guide →