activity
20232026
most citedFrom Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs

3 citations · 3 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

Ming Nie, Dan Ding, Chunwei Wang +4

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze vid…

cs.CV2025

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu +1

Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video compr…

cs.CV2025

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

Yunlong Yuan, Yuanfan Guo, Chunwei Wang +2

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power a…

cs.CV2023

Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving

Ming Nie, Renyuan Peng, Chunwei Wang +4

Large vision-language models (VLMs) have garnered increasing interest in autonomous driving areas, due to their advanced capabilities in complex reasoning tasks essential for highl…

cs.CV2023

PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection

Ming Nie, Yujing Xue, Chunwei Wang +7

Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representi…

cs.CV2023

SUIT: Learning Significance-guided Information for 3D Temporal Detection

Zheyuan Zhou, Jiachen Lu, Yihan Zeng +2

3D object detection from LiDAR point cloud is of critical importance for autonomous driving and robotics. While sequential point cloud has the potential to enhance 3D perception th…