A unified diagnostic benchmark

VLA-GAUGE:A Systematic Evaluation Framework for Generalization and Robustness of VLA Models

AuthorsTo be announced

InstitutionsTo be announced

Publication links and author information will appear here when released.

Overview

One loop, four stages

VLA-GAUGE is a unified diagnostic framework for zero-shot generalization and perturbation robustness in Vision-Language-Action models. It diagnoses failures at the capability level and turns those findings into actionable guidance for policy refinement and re-evaluation.

01Benchmark Construction
02Zero-shot Evaluation
03Robustness Evaluation
04Diagnosis-guided Refinement

Benchmark comparison

Evaluation that closes the diagnostic loop

Existing robot benchmarks typically isolate long-horizon manipulation, knowledge transfer, generalization, or robustness. VLA-GAUGE connects fine-grained capability diagnosis, ID/OOD analysis, isolated and compounded perturbations, and diagnosis-guided optimization in one evaluation loop.

6,300Diagnostic Scenarios
3,089Objects
Full support FullPartial support PartialNo support None
Comparison of VLA-GAUGE with existing robot benchmarks
BenchmarkFocusScenariosObjectsFine-grainedID/OODRobustnessOptimization
LIBEROKnowledge Transfer13075Partial supportNo supportNo supportNo support
Behavior-1KLong-horizon + Dataset1,0009,331Partial supportNo supportNo supportNo support
COLOSSEUMRobustness20147Partial supportPartial supportPartial supportPartial support
SimplerEnvGeneralization817Partial supportPartial supportPartial supportNo support
VLA-TestRobustness474Partial supportFull supportPartial supportNo support
VLABenchGeneralization + Long-horizon1002,164Partial supportPartial supportPartial supportNo support
VLA-ArenaRobustness170122Partial supportPartial supportPartial supportNo support
AGNOSTOSGeneralization23112No supportFull supportNo supportNo support
LIBERO-XRobustness60096Partial supportPartial supportPartial supportNo support
LIBERO-PRORobustness4077Partial supportFull supportPartial supportNo support
LIBERO-PlusRobustness10,030911Partial supportFull supportPartial supportPartial support
VLA-GAUGEGeneralization + Robustness6,3003,089Full supportFull supportFull supportFull support

Leaderboards

Three views of model reliability

Compare verified results across transfer, isolated sensitivity, and simultaneous perturbation stress. Leaderboards change only when you select a tab or control.

1 / 3

Zero-shot Generalization Leaderboard

Zero-Shot Score (ZS) for every model across the six capability axes: the mean performance over the ID and OOD tasks within each suite. Switch views to split each score into its in-domain and out-of-domain components, or to read the gap between them directly.

12 entries

Zero-Shot Score on each capability axis, next to the original LIBERO benchmark.

Zero-shot generalization of 12 VLA models across six capability axes, shown as Overall.
Rank
1π0.53BM — Mixed action space60.057.861.847.672.741.896.957.0
2DeepThinkVLA2.9BD — Discrete action space39.139.637.821.833.229.897.033.6
3Evo-10.77BC — Continuous action space30.228.030.78.028.223.694.824.8
4OpenVLA-OFT7BC — Continuous action space38.210.236.416.922.721.397.124.3
5UniVLA9BD — Discrete action spaceHist. — uses observation or action history29.318.723.610.235.923.695.223.6
  • 1
    π0.53BM — Mixed action space
    57.0
    Vis.
    60.0
    Sem.
    57.8
    CS-WK
    61.8
    Spa.
    47.6
    Act.
    72.7
    Comp.
    41.8
  • 2
    DeepThinkVLA2.9BD — Discrete action space
    33.6
    Vis.
    39.1
    Sem.
    39.6
    CS-WK
    37.8
    Spa.
    21.8
    Act.
    33.2
    Comp.
    29.8
  • 3
    Evo-10.77BC — Continuous action space
    24.8
    Vis.
    30.2
    Sem.
    28.0
    CS-WK
    30.7
    Spa.
    8.0
    Act.
    28.2
    Comp.
    23.6
  • 4
    OpenVLA-OFT7BC — Continuous action space
    24.3
    Vis.
    38.2
    Sem.
    10.2
    CS-WK
    36.4
    Spa.
    16.9
    Act.
    22.7
    Comp.
    21.3
  • 5
    UniVLA9BD — Discrete action spaceHist. — uses observation or action history
    23.6
    Vis.
    29.3
    Sem.
    18.7
    CS-WK
    23.6
    Spa.
    10.2
    Act.
    35.9
    Comp.
    23.6
00.0Strongest axis00.0Second strongestcompared within each row

ID and OOD scenarios are not evenly distributed across capabilities, so the overall score is a weighted combination of the two columns rather than their arithmetic mean.

Join the benchmark

Submit a model for official evaluation

Share the public Hugging Face repository containing your VLA model weights and loading instructions. The evaluation team downloads the model, runs VLA-GAUGE, and updates the verified leaderboard.

  • A readable repository. Weights must sit in a public repository. Gated repositories are accepted, but access has to be approved before the run can start.
  • Instructions we can follow. Say which checkpoint to load and list any dependency the standard setup does not cover.
  • Details for the table. Parameter count, action space, and whether the policy conditions on history — these place the model in the leaderboard groups.
  • A contact address. The team writes to you if the model fails to load or a result needs confirming.

No weights are uploaded through this website.

Submissions are emailed to the evaluation team after the repository check.

Experimental findings

From aggregate score to failure mechanism

Each analysis pairs a concise result with the underlying figure and representative rollout slots.

Key finding

High benchmark scores do not translate into reliable transfer.

Across 12 representative VLA models, zero-shot success spans only 6.6%–57.0%, while mean performance falls from 32.9% on ID scenarios to 7.2% on OOD scenarios.

6.6%–57.0%Zero-shot range
32.9% → 7.2%Mean ID → OOD
71 / 72Capability-model pairs where ID exceeds OOD
ID and OOD average success rates for each evaluated model.
The ID-to-OOD gap is consistent across nearly every model-capability pair.

Representative cases

3 cases

ID success case

Representative in-domain rollout.

OOD failure case

Representative out-of-domain rollout.

Capability contrast

Matched rollouts across capability axes.

Citation

Cite VLA-GAUGE

The official citation will be available alongside the paper release.

Reference

BibTeX

Citation metadata will be published with the VLA-GAUGE paper.

The copy action will become available when the official citation is released.