1 paper
Georgy Savva, Oscar Michel, Daohan Lu +6
Existing action-conditioned video generation models (video world models) are limited to single-agent perspectives, failing to capture the multi-agent interactions of real-world env…