22 citations · 41 across the 11 of their papers we have counts for
11 papers
Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models
Xin Wang, Wenxuan Liu, Tongtong Feng +1
Large-scale video generation models are increasingly described as world models because they can learn rich spatiotemporal regularities from visual data. However, we argue that an i…
ALAS: Adaptive Long-Horizon Action Synthesis via Async-pathway Stream Disentanglement
Yutong Shen, Hangxu Liu, Lei Zhang +4
Long-Horizon (LH) tasks in Human-Scene Interaction (HSI) are complex multi-step tasks that require continuous planning, sequential decision-making, and extended execution across do…
Self-evolving Embodied AI
Tongtong Feng, Xin Wang, Wenwu Zhu
Embodied Artificial Intelligence (AI) is an intelligent system formed by agents and their environment through active perception, embodied cognition, and action interaction. Existin…
PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
Yu-Wei Zhan, Xin Wang, Hong Chen +6
Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper und…
BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
Yu-Wei Zhan, Xin Wang, Pengzhe Mao +3
Building generalist embodied agents requires a unified system that can interpret multimodal goals, model environment dynamics, and execute reliable actions across diverse real-worl…
Embodied AI: From LLMs to World Models
Tongtong Feng, Xin Wang, Yu-Gang Jiang +1
Embodied Artificial Intelligence (AI) is an intelligent system paradigm for achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications and…