3 papers
cs.CV2026
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
Haichen He, Jiayi Zhou, Sifeng Shang +3
Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video du…
cs.CV2024
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
Yabin Zhang, Wenjie Zhu, Hui Tang +3
With the emergence of pre-trained vision-language models like CLIP, how to adapt them to various downstream classification tasks has garnered significant attention in recent resear…
cs.LG2024
Open-Vocabulary Calibration for Fine-tuned CLIP
Shuoyuan Wang, Jindong Wang, Guoqing Wang +3
Vision-language models (VLMs) have emerged as formidable tools, showing their strong capability in handling various open-vocabulary tasks in image recognition, text-driven visual c…