15 papers
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
Siyu Xu, Yunke Wang, Zijian Wang +6
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
Bingzhou Li, Tao Huang
Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Exist…
Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation
Fengnian Zhang, Tao Huang, Siyu Xu +2
Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs).…
Differentiable Efficient Operator Search
Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4
Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…
Retrieval-Augmented Linguistic Calibration
Yi-Fan Yeh, Linwei Tao, Minjing Dong +4
Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic c…
Adversarial Error Correction for Visual Autoregressive Generation
Ligong Bi, Tao Huang, Jianyuan Guo +1
Visual Autoregressive (VAR) models have emerged as a powerful paradigm for image synthesis by performing hierarchical next-scale prediction. However, VAR models are inherently pron…