2 papers
cs.CV2025
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Yuxuan Wang, Yiqi Song, Cihang Xie +2
Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demand…
cs.CV2024
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
Yayun Qi, Hongxi Li, Yiqi Song +2
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelli…