7 papers
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Chonghuinan Wang, Zhikai Chen, Chunwei Wang +9
The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving…
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
Runhui Huang, Xinpeng Ding, Chunwei Wang +7
High-resolution inputs enable Large Vision-Language Models (LVLMs) to discern finer visual details, enhancing their comprehension capabilities. To reduce the training and computati…
Translating Images to Road Network: A Sequence-to-Sequence Perspective
Jiachen Lu, Ming Nie, Bozhou Zhang +6
The extraction of road network is essential for the generation of high-definition maps since it enables the precise localization of road landmarks and their interconnections. Howev…
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
Runhui Huang, Chunwei Wang, Junwei Yang +8
We present ILLUME+ that leverages dual visual tokenization and a diffusion decoder to improve both deep semantic understanding and high-fidelity image generation. Existing unified…
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
Xinpeng Ding, Jianhua Han, Hang Xu +2
Recent efforts to use natural language for interpretable driving focus mainly on planning, neglecting perception tasks. In this paper, we address this gap by introducing ROLISP (Ri…
FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise
Yunlong Yuan, Yuanfan Guo, Chunwei Wang +3
Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise pri…