From the 1 of 4 linked papers with an AI index.
4 papers
MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
Ke Cheng, Hanqiao Ye, Lei Shi +8
The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
Lei Shi, Andreas Bulling
We propose CLAD, a Constrained Latent Action Diffusion model for vision-language procedure planning in instructional videos. Procedure planning is the challenging task of predictin…
VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning
Li-Heng Chen, Ke Cheng, Yahui Liu +3
Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi +1
Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Un…