3 papers
cs.CV2026
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
Ming Nie, Dan Ding, Chunwei Wang +4
Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze vid…
cs.CV2025
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
Ming Nie, Chunwei Wang, Hang Xu +1
Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video compr…
cs.CV2025
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
Yunlong Yuan, Yuanfan Guo, Chunwei Wang +2
Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power a…