5 papers
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
Ming Nie, Chunwei Wang, Jianhua Han +2
Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, w…
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
Ming Nie, Dan Ding, Chunwei Wang +4
Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze vid…
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
Ming Nie, Chunwei Wang, Hang Xu +1
Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video compr…
CipherMind: The Longest Codebook in the World
Ming Nie, Zhixiong Yang, Bingsheng Wei
In recent years, the widespread application of large language models has inspired us to consider using inference for communication encryption. We therefore propose CipherMind, whic…
LaneCorrect: Self-supervised Lane Detection
Ming Nie, Xinyue Cai, Hang Xu +1
Lane detection has evolved highly functional autonomous driving system to understand driving scenes even under complex environments. In this paper, we work towards developing a gen…