5 papers
DIMOS: Disentangling Instance-level Moving Object Segmentation
Hongxiang Huang, Hongwei Ren, Xiaopeng Lin +3
Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras recor…
Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models
Zhiming Liu, Yujie Wei, Lei Feng +5
Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downst…
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
Yujie Wei, Jiahan Fan, Jiyu Guo +6
Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a c…
UtilGen: Utility-Centric Generative Data Augmentation with Dual-Level Task Adaptation
Jiyu Guo, Shuo Yang, Yiming Huang +6
Data augmentation using generative models has emerged as a powerful paradigm for enhancing performance in computer vision tasks. However, most existing augmentation approaches prim…
Understanding Data Influence with Differential Approximation
Haoru Tan, Sitong Wu, Xiuzhe Wu +5
Data plays a pivotal role in the groundbreaking advancements in artificial intelligence. The quantitative analysis of data significantly contributes to model training, enhancing bo…