works on

From the 1 of 26 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.ROShow all

6 papers · 1 filter

cs.RO2026

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

Zijian Song, Qichang Li, Sihan Qin +4

The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning…

cs.RO2026

Stable Language Guidance for Vision-Language-Action Models

Zhihao Zhan, Yuhao Chen, Jiaying Zhou +5

Vision-Language-Action (VLA) models have demonstrated impressive capabilities in generalized robotic control; however, they remain notoriously brittle to linguistic perturbations.…

cs.RO2026

Robotic Manipulation is Vision-to-Geometry Mapping: Vision-Geometry Backbones over Language and Video Models

Zijian Song, Qichang Li, Jiawei Zhou +4

At its core, robotic manipulation is a problem of vision-to-geometry mapping (). Physical actions are fundamentally defined by geometric properties like 3D posi…

cs.RO2026

VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

Weiqi Li, Quande Zhang, Ruifeng Zhai +2

Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittle…

cs.RO2026

E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusion

Zhihao Zhan, Jiaying Zhou, Likui Zhang +10

Vision-Language-Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, ex…

cs.RO2026

RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation

Yuhao Chen, Zhihao Zhan, Xiaoxin Lin +11

VLA models have achieved remarkable progress in embodied intelligence; however, their evaluation remains largely confined to simulations or highly constrained real-world settings.…