3 papers
cs.CV2026
PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
Zihan Song, Shuo Ye, Bo Zhao +4
Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-duration frames remains a hindranc…
cs.CV2026
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Yixin Ji, Fanghua Ye, Juntao Li +5
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing…
cs.CV2026
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
Quoc-Huy Trinh, Mustapha Abdullahi, Bo Zhao +1
Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment i…