5 papers
DriveFix: Spatio-Temporally Coherent Driving Scene Restoration
Heyu Si, Brandon James Denis, Muyang Sun +9
Recent advancements in 4D scene reconstruction, particularly those leveraging diffusion priors, have shown promise for novel view synthesis in autonomous driving. However, these me…
FAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving
Yaoru Li, Federico Landi, Marco Godi +6
Despite rapid progress in autonomous driving, reliable training and evaluation of driving systems remain fundamentally constrained by the lack of scalable and interactive simulatio…
Parallelized Planning-Acting for Efficient LLM-based Multi-Agent Systems in Minecraft
Yaoru Li, Shunyu Liu, Tongya Zheng +2
Recent advancements in Large Language Model~(LLM)-based Multi-Agent Systems (MAS) have demonstrated remarkable potential for tackling complex decision-making tasks. However, existi…
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
Yaoru Li, Heyu Si, Federico Landi +8
Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing…
Odyssey: Empowering Minecraft Agents with Open-World Skills
Shunyu Liu, Yaoru Li, Kongcheng Zhang +5
Recent studies have delved into constructing generalist agents for open-world environments like Minecraft. Despite the encouraging results, existing efforts mainly focus on solving…