9 papers
Calibrated Multimodal Representation Learning with Missing Modalities
Xiaohao Liu, Xiaobo Xia, Jiaheng Wei +4
Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditional cross-modal alignment to pro…
FreeAct: Freeing Activations for LLM Quantization
Xiaohao Liu, Xiaobo Xia, Manyi Zhang +6
Quantization is pivotal for mitigating the significant memory and computational overhead of Large Language Models (LLMs). While emerging transformation-based methods have successfu…
Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models
Zhiming Liu, Yujie Wei, Lei Feng +5
Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downst…
Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
Yujie Wei, Jiahan Fan, Jiyu Guo +5
Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a c…
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
Weisheng Dai, Kai Lan, Jianyi Zhou +5
Vision-Language-Action (VLA) models achieve preliminary generalization through pretraining on large scale robot teleoperation datasets. However, acquiring datasets that comprehensi…
APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation
Daoxuan Zhang, Ping Chen, Xiaobo Xia +4
Aerial Object Goal Navigation, a challenging frontier in Embodied AI, requires an Unmanned Aerial Vehicle (UAV) agent to autonomously explore, reason, and identify a specific targe…