2 papers
cs.CV2024
Discriminative Spatial-Semantic VOS Solution: 1st Place Solution for 6th LSVOS
Deshui Miao, Yameng Gu, Xin Li +3
Video object segmentation (VOS) is a crucial task in computer vision, but current VOS methods struggle with complex scenes and prolonged object motions. To address these challenges…
cs.CL2024
LLAVADI: What Matters For Multimodal Large Language Models Distillation
Shilin Xu, Xiangtai Li, Haobo Yuan +3
The recent surge in Multimodal Large Language Models (MLLMs) has showcased their remarkable potential for achieving generalized intelligence by integrating visual understanding int…