A Causality-aware Infer-Diagnose-Refine Framework
for Test-time Modality Adaptation in VLA Models

Haoyu Zhang1 Yuwei Wu1 Jin Chen2† Gao Zhi1†
Zhenxin Diao1 Mingyang Gao1 Kun Wu3 Yongchun Liu2 Fan Li2
1Beijing Institute of Technology    2AInnovation Co. Ltd.    3Beijing Innovation Center of Humanoid Robotics
Corresponding authors
arXiv Code

Abstract

IDR Framework Overview

Vision-language-action (VLA) models predict sequential actions to execute tasks specified by language instructions, conditioned on visual observations and proprioceptive states. However, how to fuse modalities in VLA models remains an open problem, since robot manipulation involves dynamic phases, such as long-distance movements and close-range interactions, in which the importance of visual observations may vary over time.

In this paper, we propose an Infer-Diagnose-Refine (IDR) framework, a model-agnostic framework that can be integrated with diverse VLA architectures for refining action predictions at test time. IDR first infers actions under factual and counterfactual scenarios of visual observations, and then diagnoses the causal effects of visual observations as the estimated dynamic importance, which is finally used to refine the action predictions in a training-free manner. We further design a causality-aware action refiner to realize the IDR framework, including zero-padding interventions for inferring counterfactual actions, norm-based quantification for diagnosing causal effects, and gated residual fusion for refining actions. Extensive experiments on both simulation benchmarks and real-world tasks show improvements in overall performance across multiple VLA backbones, demonstrating the efficacy of dynamically adjusting visual importance at test time.

Method Overview

IDR Method

IDR Framework Architecture: Zero-padding interventions for counterfactual inference, norm-based effect quantification, and gated residual fusion.

Key Components

Infer: IDR constructs counterfactual inputs through zero-padding interventions, enabling the frozen VLA model to infer factual and counterfactual outputs under different modality conditions.

Diagnose: The changes of outputs are quantified with a norm-based effect measurement to diagnose the dynamic importance of visual observations.

Refine: The diagnosed effects are integrated with the initial action prediction through gated residual fusion, producing a refined action while preserving low-level control stability.

Causal Analysis

Causal Graph

Causal graph illustrating the modality intervention framework.

IDR treats visual importance as a dynamic factor during execution, diagnoses how visual observations causally affect the current prediction, and refines the action accordingly.

Observations from Test-Time Modality Diagnosis

Through test-time causal inference across models and environments, we derive three key observations. To compare visual effects across models and action scales, we compute the normalized visual effect ratio Rimg = Eimg / (Eimg + Eprop + ε).

Observation 1: Deployment environments shift inherent causal effect patterns

Environment Suite/Task Eimg Eprop Rimg
LIBEROAll suites0.671.5929.7%
WidowXReal-robot tasks0.620.7644.9%
CALVINABC→D0.610.6050.7%
SIMPLERGoogle Variant Aggregation1.180.0993.2%
SIMPLERGoogle Visual Matching0.990.0793.8%

The same VLA model (X-VLA) shifts its visual effect pattern substantially across different benchmarks or embodiments. X-VLA is proprioception-primary in LIBERO but becomes vision-dominant in SIMPLER. This environment-dependent shift suggests that modality balance is also associated with deployment environments.

Observation 2: Model architectures shape specific causal effect patterns

Model Suite Heatmap

Rimg heatmap across models and LIBERO suites. Each model maintains a consistent visual effect pattern.

Within the same benchmark, different VLA models exhibit significantly different causal effect patterns. This variation reflects architectural designs where models relying heavily on a dominant vision-language space maintain a higher Rimg, while architectures that inject proprioceptive states and action tokens earlier show stronger proprioceptive effects. As shown in the figure, π0.5 and OpenVLA-OFT are vision-heavy, X-VLA is proprioception-primary, and VLA-Adapter is vision-dominant.

Observation 3: Manipulation phases dynamically alter causal effect patterns

Phase-aligned Effect Ratio

Phase-aligned visual effect ratio on π0.5. The plot reports the visual effect ratio Rimg over task progress for the baseline, Mode E, and Mode F. Dashed vertical lines indicate gripper close and open events.

Within a single task execution, modality effects change with manipulation phases. The baseline Rimg fluctuates with task progress and changes around gripper close/open events, highlighting the need for dynamic adaptation during execution. Visual effect tends to increase during localization and placement, where external scene information is critical, while proprioceptive effect becomes more influential during motion-transition phases.

Simulation Benchmark Results

LIBERO Benchmark

Size Method Spatial Object Goal Long Avg
Large (≥4B)OpenVLA-OFT92.6099.2096.2094.6095.65
+IDR94.2099.2098.0094.8096.55
Small (<4B)π0.597.5095.0097.0095.5096.25
+IDR98.00100.0097.0095.0097.50
Tiny (<1B)X-VLA92.6099.2098.0094.4096.05
+IDR95.00100.0093.0093.0096.50
VLA-Adapter94.2095.6098.0092.6095.10
+IDR97.00100.0096.0094.0096.50

IDR consistently improves all four VLA backbones. The largest gains are observed on VLA-Adapter (+1.40) and π0.5 (+1.25), where the visual-effect-guided adaptation provides the most benefit.

CALVIN (ABC→D) Benchmark

Method 1/5 2/5 3/5 4/5 5/5 Avg Len
X-VLA95.790.986.078.870.54.22
+IDR98.194.289.984.377.34.44
VLA-Adapter97.193.888.882.576.44.39
+IDR97.794.189.383.877.84.43

On CALVIN, IDR improves both the average completed length (from 4.22 to 4.44) and the 5/5 completion rate (from 70.5% to 77.3%), demonstrating its effectiveness in long-horizon execution.

SIMPLER Benchmark (on X-VLA)

Environment Method Blocks Eggplant Spoon Place Avg
Google VMX-VLA98.3397.0871.7657.4181.15
+IDR96.6796.6775.9259.2682.13
Google VAX-VLA86.2583.8360.9559.2675.88
+IDR89.6383.8068.1074.0778.90
WidowXX-VLA75.00100.00100.00100.0093.75
+IDR87.50100.00100.0095.8095.83

On SIMPLER, the largest gain is achieved in the Variant Aggregation (VA) setting (+3.02), where visually grounded manipulation is particularly critical.

LIBERO Benchmark Comparison

Comparison of baseline (left) vs IDR (right) on LIBERO tasks. Each pair shows failure case without IDR and success case with IDR.

π0.5

π05 Baseline Failed

Baseline - Failed

π05 IDR Successful

+IDR - Successful

Put both moka pots on the stove

OpenVLA-OFT

OpenVLA Baseline Failed

Baseline - Failed

OpenVLA IDR Successful

+IDR - Successful

Put both the alphabet soup and the tomato sauce in the basket

VLA-Adapter

VLA-Adapter Baseline Failed

Baseline - Failed

VLA-Adapter IDR Successful

+IDR - Successful

Put the yellow mug in the microwave

X-VLA

X-VLA Baseline Failed

Baseline - Failed

X-VLA IDR Successful

+IDR - Successful

Turn on the stove and put the moka pot on it

SIMPLER Benchmark Comparison

Comparison of X-VLA (left) vs X-VLA+IDR (right) on SIMPLER tasks across different embodiments.

Google Robot (VA)

Google VA Baseline Failed

Baseline - Failed

Google VA IDR Successful

+IDR - Successful

Open top drawer and place coke_can into top drawer

Google Robot (VM)

Google VM Baseline Failed

Baseline - Failed

Google VM IDR Successful

+IDR - Successful

Open and close bottom drawer

WidowX

WidowX Baseline Failed

Baseline - Failed

WidowX IDR Successful

+IDR - Successful

Put eggplant into yellow basket

Real-World Experiment Results

We evaluate IDR on real-world manipulation tasks using a dual-arm ARX5 platform. The evaluation tasks cover four task categories with different manipulation challenges.

Real World Results

Real-world experiment results on ARX5 dual-arm platform.

Performance on Fold Clothes, Grab Cola and Sweep Trash

Task Level Method Succ. (%) Avg. Time
Grab Cola Level 1 Baseline 95.8 8.4s
Ours 100.0 8.7s
Level 2 Baseline 66.7 14.3s
Ours 91.7 13.4s
Level 3 Baseline 91.7 21.3s
Ours 91.7 19.2s
Fold Clothes Level 1 Baseline 66.7 47s
Ours 91.7 37s
Level 2 Baseline 83.3 52s
Ours 91.7 47s
Level 3 Baseline 8.3 222s
Ours 33.3 155s
Sweep Trash Baseline 59.4 16.0s
Ours 68.8 15.9s

Organize Table: Long-horizon Sequential Task

Tasks Completed in a Row
Method 1/5 2/5 3/5 4/5 5/5 Avg Len ↑
Baseline83.3%66.7%33.3%33.3%0.0%2.167
Ours100.0%100.0%100.0%66.7%33.3%4.000

On the challenging 5-step Organize Table task, IDR achieves a 33.3% full-completion rate compared with 0% for the baseline, while increasing the average completed steps from 2.17 to 4.00 out of 5.

Overall, IDR improves manipulation performance across various real-world tasks, particularly on challenging tasks requiring fine-grained visual guidance and contact-rich manipulation.

Ablation Studies

Hyperparameter Sensitivity

Hyperparameter Ablation

Moderate correction scales and intervention thresholds yield the best performance, while excessive correction or uniform intervention degrades action generation.

Correction scale (α): Performance peaks at α=0.08, while nearby values such as α=0.05, 0.10 produce comparable results. Larger values (α ≥ 0.5) degrade performance, indicating that excessive visual-effect correction can destabilize action generation.

Intervention threshold (τ): Small values rarely activate the gate, and therefore provide limited improvement. Uniform intervention with τ=999 performs below the baseline. The best performance is obtained at τ=7, where refinement is applied only to predictions with low diagnosed visual effect.

Real-World Benchmark Comparison

Side-by-side comparison of baseline (left) and IDR (right) on real-world tasks using a dual-arm ARX5 platform. Videos are shown at 2× speed.

Fold Clothes

Baseline - Failed

+IDR - Successful

Deformable object manipulation — fold clothes

Sweep Trash

Baseline - Failed

+IDR - Successful

Contact-rich sweeping — sweep trash into a dustpan

Grab Cola

Baseline - Failed

+IDR - Successful

Bimanual coordination — grab and transfer a cola can

Organize Table

Baseline - Failed

+IDR - Successful

Long-horizon sequential — place 5 objects into a drawer

BibTeX

@article{zhang2026causality,
  title={A Causality-aware Infer-diagnose-refine Framework for Test-time Modality Adaptation in VLA Models},
  author={Zhang, Haoyu and Wu, Yuwei and Chen, Jin and Zhi, Gao and Diao, Zhenxin and Gao, Mingyang and Wu, Kun and Liu, Yongchun and Li, Fan},
  journal={arXiv preprint arXiv:2607.25516},
  year={2026}
}