activity
20242026
collaborators

15 papers

cs.CV2026

LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration

Enhuai Liu, Yunke Wang, Yutong Wang +2

Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic cost that dominates as resolution and dura…

cs.RO2026

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

Siyu Xu, Yunke Wang, Zijian Wang +6

Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…

cs.RO2026

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

Fengnian Zhang, Tao Huang, Siyu Xu +2

Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs).…

cs.LG2026

Differentiable Efficient Operator Search

Xiaohuan Pei, Jiyuan Zhang, Yuanfan Guo +4

Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operat…

cs.RO2026

See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model

Yixu Feng, Zinan Zhao, Yanxiang Ma +4

Vision-Language-Action (VLA) models have shown remarkable promise in robotics manipulation, yet their high computational cost hinders real-time deployment. Existing token pruning m…

cs.CV2026

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation

Jingxuan He, Xiyu Wang, Yunke Wang +2

Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which impl…