works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.RO2026

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

Jiayi Luo, Hanxin Zhu, Chen Gao +5

Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable p…

cs.RO2026

TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction

Lei Jin, Yiding Ma, Xin Zhang +3

The paper introduces TacWAM, a mechanics-aware tactile world action model that predicts future tactile signals and uses them as supervision for training contact-rich robot manipula…

cs.CV2026

CP4D: Compositional Physics-aware 4D Scene Generation

Hanxin Zhu, Cong Wang, Tianyu He +4

4D generation (\textit{i.e.}, dynamic 3D generation) has recently emerged as a rapidly growing research frontier due to its powerful spatiotemporal modeling capabilities. However,…

cs.CV2026

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation

Jiayi Luo, Qiyan Liu, Tengyang Wang +8

Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens.…

cs.CV2026

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

Wei Luo, Yiting Lu, Xin Li +32

This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The…

cs.CV2026

Semantic Audio-Visual Navigation in Continuous Environments

Yichen Zeng, Hebaixu Wang, Meng Liu +4

Audio-visual navigation enables embodied agents to navigate toward sound-emitting targets by leveraging both auditory and visual cues. However, most existing approaches rely on pre…