382 citations · 389 across the 13 of their papers we have counts for
Showing 2026Show all
3 papers · 1 filter
cs.RO2026
Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models
Xingyu Ding, Yuzhong Zhao, Chunhai Zhao +3
Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with lo…
cs.RO2026
Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
Xingyu Ding, Yuzhong Zhao, Yang Wu +4
Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-ag…
cs.CL2026
Balancing Understanding and Generation in Discrete Diffusion Models
Yue Liu, Yuzhong Zhao, Zheyong Xie +5
In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot ge…