papers

Publications (9)

cs.RO2026

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Junhao Shi, Siyin Wang, Xiaopeng Yu +3

Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly t…

cs.RO2024

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Shiduo Zhang, Zhe Xu, Peiju Liu +8

General-purposed embodied agents are designed to understand the users' natural instructions or intentions and act precisely to complete universal tasks. Recently, methods based on…

cs.CV2025

FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization

Yicheng Liu, Shiduo Zhang, Zibin Dong +12

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often…

cs.LG2022

Model-Based Opponent Modeling

Xiaopeng Yu, Jiechuan Jiang, Wanpeng Zhang +2

When one agent interacts with a multi-agent environment, it is challenging to deal with various opponents unseen before. Modeling the behaviors, goals, or beliefs of opponents coul…

cs.RO2026

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

Senyu Fei, Xiaopeng Yu, Siyin Wang +3

Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods, critic-based approaches rely o…

physics.optics2020

Time-reversed photoacoustic guided time-reversed ultrasonically encoded optical focusing

Juze Zhang, Zijian Gao, Xiaopeng Yu +3

Deep-tissue optical imaging is a longstanding challenge limited by scattering. Both optical imaging and treatment can benefit from focusing light in deep tissue beyond one transpor…

cs.RO2026

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

Li Ji, Siyin Wang, Pengfang Qian +5

Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance…

cs.RO2026

Coarse-to-Control: Action-Token Planning for Vision-Language-Action Models

Jinhao Wu, Shiduo Zhang, Yicheng Liu +9

Most vision-language-action (VLA) models map observations directly to actions without explicit intermediate planning, which limits performance on long-horizon tasks where early mis…

cs.DC2022

Predicting the Output Structure of Sparse Matrix Multiplication with Sampled Compression Ratio

Zhaoyang Du, Yijin Guan, Tianchan Guan +7

Sparse general matrix multiplication (SpGEMM) is a fundamental building block in numerous scientific applications. One critical task of SpGEMM is to compute or predict the structur…