2 papers
cs.CV2025
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
Zhao Wang, Aoxue Li, Lingting Zhu +3
Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generatio…
cs.CV2024
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
Zhenhua Xu, Yujia Zhang, Enze Xie +5
Multimodal large language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-text…