1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2024
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
Yayun Qi, Hongxi Li, Yiqi Song +2
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelli…
cs.CV2024★ 1 cited
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
Yuxuan Wang, Yiqi Song, Cihang Xie +2
Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demand…