56 citations · 56 across the 3 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
M-LLM Based Video Frame Selection for Efficient Video Understanding
Kai Hu, Feng Gao, Xiaohan Nie +8
Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply n…
cs.CV2024
Multimodal Instruction Tuning with Hybrid State Space Models
Jianing Zhou, Han Li, Shuai Zhang +5
Handling lengthy context is crucial for enhancing the recognition and understanding capabilities of multimodal large language models (MLLMs) in applications such as processing high…
cs.CV2014★ 56 cited
Cross-view Action Modeling, Learning and Recognition
Jiang wang, Xiaohan Nie, Yin Xia +2
Existing methods on video-based action recognition are generally view-dependent, i.e., performing recognition from the same views seen in the training data. We present a novel mult…