From the 1 of 141 linked papers with an AI index.
1 citations · 1 across the 50 of their papers we have counts for
89 papers · 1 filter
A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
Jingtao Sun, Xiaohai He, Yike Zhang +4
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a…
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
Quanjian Song, Yiren Song, Kelly Peng +2
WorldWander is a framework that translates video content between first‑person (egocentric) and third‑person (exocentric) views using video diffusion transformers and in‑context lea…
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Wenzheng Zeng, Siyi Jiao, Chen Gao +2
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation. In this domain, autoregres…
Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
Rui Zhao, Kaiming Yang, Jifeng Zhu +6
Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follow…
Demo2Tutorial: From Human Experience to Multimodal Software Tutorials
Zechen Bai, Zhiheng Chen, Yiqi Lin +5
Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowledge. We introduce Demo2Tutori…
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
Heyuan Gao, Bangxun Tang, Yiren Song +4
We present PAI-Studio, a new reference-conditioned video synthesis task that addresses a long-standing challenge in cinematic background replacement: generating dynamic backgrounds…