Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
M-LLM Based Video Frame Selection for Efficient Video Understanding
Kai Hu, Feng Gao, Xiaohan Nie +8
Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply n…
cs.CV2024
MV2MAE: Multi-View Video Masked Autoencoders
Ketul Shah, Robert Crandall, Jie Xu +4
Videos captured from multiple viewpoints can help in perceiving the 3D structure of the world and benefit computer vision tasks such as action recognition, tracking, etc. In this p…