Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
InstructionBench: An Instructional Video Understanding Benchmark
Haiwan Wei, Yitian Yuan, Xiaohan Lan +2
Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insuffic…
cs.CV2024
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Shimin Chen, Xiaohan Lan, Yitian Yuan +2
Rapid development of large language models (LLMs) has significantly advanced multimodal large language models (LMMs), particularly in vision-language tasks. However, existing video…
cs.CV2024
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
Xiaohan Lan, Yitian Yuan, Zequn Jie +1
Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of…