10 papers
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
Zihan Song, Shuo Ye, Bo Zhao +4
Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
Jiayu Zhang, Shuo Ye, Qilang Ye +3
Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes.…
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
Yelin Wang, Zijia Song, Shuo Ye +6
Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and application value. However, mos…
DeceptionX: From Multimodal Evidence to Explainable Deception Detection
Jiayu Zhang, Shuo Ye, Jiajian Huang +8
Deception detection is a critical and highly challenging task within affective computing and behavioral analysis. Existing deep learning methods typically treat this task as a stra…
Seg-Agent: Test-Time Multimodal Reasoning for Training-Free Language-Guided Segmentation
Chao Hao, Jun Xu, Ji Du +6
Language-guided segmentation transcends the scope limitations of traditional semantic segmentation, enabling models to segment arbitrary target regions based on natural language in…
Text-Guided Multimodal Unified Industrial Anomaly Detection
Zewen Li, Shuo Ye, Zitong Yu +2
Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing unsupervised methods suffer…