2 citations · 3 across the 4 of their papers we have counts for
5 papers
Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
Xiangbo Gao, Siyuan Yang, Ping He +12
We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generation, and the update becomes visib…
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
Shuo Xing, Lanqing Guo, Hongyuan Hua +5
Recent Multimodal Large Language Models (MLLMs) excel on benchmark vision-language tasks, yet little is known about how input visual quality shapes their responses. Does higher per…
Generative AI for Autonomous Driving: Frontiers and Opportunities
Yuping Wang, Shuo Xing, Cui Can +44
Generative Artificial Intelligence (GenAI) constitutes a transformative technological wave that reconfigures industries through its unparalleled capabilities for content creation,…
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
Shuo Xing, Chengyuan Qian, Yuping Wang +4
Since the advent of Multimodal Large Language Models (MLLMs), they have made a significant impact across a wide range of real-world applications, particularly in Autonomous Driving…
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
Shuo Xing, Hongyuan Hua, Xiangbo Gao +10
Recent advancements in large vision language models (VLMs) tailored for autonomous driving (AD) have shown strong scene understanding and reasoning capabilities, making them undeni…