3 papers
cs.CV2026
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
Zhaowei Wang, Lishu Luo, Haodong Duan +9
Long-context modeling is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-document understanding, video…
cs.CV2024
Streaming Dense Video Captioning
Xingyi Zhou, Anurag Arnab, Shyamal Buch +5
An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descr…
cs.CV2023
Pixel Aligned Language Models
Jiarui Xu, Xingyi Zhou, Shen Yan +5
Large language models have achieved great success in recent years, so as their variants in vision. Existing vision-language models can describe images in natural languages, answer…