15 papers
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Enhuai Liu, Yunke Wang, Yutong Wang +2
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and dura…
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
Siyu Xu, Yunke Wang, Zijian Wang +6
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…
Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation
Fengnian Zhang, Tao Huang, Siyu Xu +2
Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs).…
Differentiable Efficient Operator Search
Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4
Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
Yixu Feng, Zinan Zhao, Yanxiang Ma +4
Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning m…
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
Jingxuan He, Xiyu Wang, Yunke Wang +2
Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which impl…