activity
20242026
collaborators

15 papers

cs.RO2026

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

Siyu Xu, Yunke Wang, Zijian Wang +6

Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…

cs.MM2026

DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression

Bingzhou Li, Tao Huang

Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Exist…

cs.RO2026

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

Fengnian Zhang, Tao Huang, Siyu Xu +2

Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs).…

cs.LG2026

Differentiable Efficient Operator Search

Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4

Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…

cs.CL2026

Retrieval-Augmented Linguistic Calibration

Yi-Fan Yeh, Linwei Tao, Minjing Dong +4

Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic c…

cs.CV2026

Adversarial Error Correction for Visual Autoregressive Generation

Ligong Bi, Tao Huang, Jianyuan Guo +1

Visual Autoregressive (VAR) models have emerged as a powerful paradigm for image synthesis by performing hierarchical next-scale prediction. However, VAR models are inherently pron…