6 citations · 10 across the 26 of their papers we have counts for
Showing 2024 · cs.CVShow all
2 papers · 2 filters
cs.CV2024
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
Jungang Li, Sicheng Tao, Yibo Yan +6
Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs…
cs.CV2024
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
Xu Zheng, Haiwei Xue, Jialei Chen +6
Simultaneously using multimodal inputs from multiple sensors to train segmentors is intuitively advantageous but practically challenging. A key challenge is unimodal bias, where mu…