From the 1 of 4 linked papers with an AI index.
9 papers
MWorld: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming
Ke Cheng, Hanqiao Ye, Lei Shi +8
The paper introduces M⁴World, a multimodal driving world model that generates synchronized surround-view video and LiDAR streams while allowing fine-grained, interactive manipulati…
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
Lei Shi, Andreas Bulling
We propose CLAD, a Constrained Latent Action Diffusion model for vision-language procedure planning in instructional videos. Procedure planning is the challenging task of predictin…
VistaGEN: Consistent Driving Video Generation with Fine-Grained Control Using Multiview Visual-Language Reasoning
Li-Heng Chen, Ke Cheng, Yahui Liu +3
Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse dri…
Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi +1
Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Un…
Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social Interactions
Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi +1
We propose MToMnet - a Theory of Mind (ToM) neural network for predicting beliefs and their dynamics during human social interactions from multimodal input. ToM is key for effectiv…
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
Lei Shi, Paul Bürkner, Andreas Bulling
We present ActionDiffusion -- a novel diffusion model for procedure planning in instructional videos that is the first to take temporal inter-dependencies between actions into acco…