From the 1 of 10 linked papers with an AI index.
10 papers
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
Jiayi Luo, Hanxin Zhu, Chen Gao +5
Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…
TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction
Lei Jin, Yiding Ma, Xin Zhang +3
The paper introduces TacWAM, a mechanics-aware tactile world action model that predicts future tactile signals and uses them as supervision for training contact-rich robot manipula…
CP4D: Compositional Physics-aware 4D Scene Generation
Hanxin Zhu, Cong Wang, Tianyu He +4
4D generation (\textit{i.e.}, dynamic 3D generation) has recently emerged as a rapidly growing research frontier due to its powerful spatiotemporal modeling capabilities. However,…
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
Jiayi Luo, Qiyan Liu, Tengyang Wang +8
Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens.…
LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)
Wei Luo, Yiting Lu, Xin Li +32
This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The…
Semantic Audio-Visual Navigation in Continuous Environments
Yichen Zeng, Hebaixu Wang, Meng Liu +4
Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on pre…