collaborators

8 papers

cs.CV2026

Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization

Ming Nie, Chunwei Wang, Jianhua Han +2

Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, w…

cs.CV2026

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

Ming Nie, Dan Ding, Chunwei Wang +4

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze vid…

cs.CV2026

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

Jiahui Zhang, Yurui Chen, Yanpeng Zhou +10

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unl…

cs.CV2025

Translating Images to Road Network: A Sequence-to-Sequence Perspective

Jiachen Lu, Ming Nie, Bozhou Zhang +6

The extraction of road network is essential for the generation of high-definition maps since it enables the precise localization of road landmarks and their interconnections. Howev…

cs.CV2025

KFFocus: Highlighting Keyframes for Enhanced Video Understanding

Ming Nie, Chunwei Wang, Hang Xu +1

Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video compr…

cs.CV2025

LaneCorrect: Self-supervised Lane Detection

Ming Nie, Xinyue Cai, Hang Xu +1

Lane detection has evolved highly functional autonomous driving system to understand driving scenes even under complex environments. In this paper, we work towards developing a gen…