4 papers
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
Hui Geng, Yi Su, Han Yin +7
Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To adapt these models to downstream tasks w…
Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection
Wenhao Li, Zimeng Wu, Yu Wu +2
Unmanned aerial vehicle (UAV) based object detection is a critical but challenging task, when applied in dynamically changing scenarios with limited annotated training data. Layout…
Collaborative Multi-Mode Pruning for Vision-Language Models
Zimeng Wu, Yunhong Wang, Donghao Wang +1
Vision-Language Models (VLMs) have advanced rapidly within the unified Transformer architecture, yet their deployment on resource-constrained devices remains challenging due to hig…
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
Ruizhe Ou, Yuan Hu, Fan Zhang +2
Multi-modal large language models (MLLMs) have achieved remarkable success in image- and region-level remote sensing (RS) image understanding tasks, such as image captioning, visua…