collaborators

5 papers

cs.CV2026

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

Ming Nie, Dan Ding, Chunwei Wang +4

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze vid…

cs.CV2025

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu +1

Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video compr…

cs.CV2025

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

Jiahui Zhang, Yurui Chen, Yanpeng Zhou +10

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unl…

eess.IV2025

FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise

Yunlong Yuan, Yuanfan Guo, Chunwei Wang +3

Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise pri…

cs.CV2025

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

Yunlong Yuan, Yuanfan Guo, Chunwei Wang +2

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power a…