3 citations · 3 across the 4 of their papers we have counts for
6 papers · 1 filter
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
Ming Nie, Dan Ding, Chunwei Wang +4
Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze vid…
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
Ming Nie, Chunwei Wang, Hang Xu +1
Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video compr…
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
Yunlong Yuan, Yuanfan Guo, Chunwei Wang +2
Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power a…
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
Ming Nie, Renyuan Peng, Chunwei Wang +4
Large vision-language models (VLMs) have garnered increasing interest in autonomous driving areas, due to their advanced capabilities in complex reasoning tasks essential for highl…
PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection
Ming Nie, Yujing Xue, Chunwei Wang +7
Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representi…
SUIT: Learning Significance-guided Information for 3D Temporal Detection
Zheyuan Zhou, Jiachen Lu, Yihan Zeng +2
3D object detection from LiDAR point cloud is of critical importance for autonomous driving and robotics. While sequential point cloud has the potential to enhance 3D perception th…