4 papers
UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
Hao Lu, Ziyang Liu, Guangfeng Jiang +4
Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods…
ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
Hao Lu, Jiahao Wang, Yaolun Zhang +5
Video multimodal large language models (Video-MLLMs) have achieved remarkable progress in video understanding. However, they remain vulnerable to hallucination-producing content in…
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Yujia Liang, Jile Jiao, Xuetao Feng +3
Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with vary…
PairEdit: Learning Semantic Variations for Exemplar-based Image Editing
Haoguang Lu, Jiacheng Chen, Zhenguo Yang +4
Recent advancements in text-guided image editing have achieved notable success by leveraging natural language prompts for fine-grained semantic control. However, certain editing se…