works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

Autonomous Video Generation with Counterfactual Controllability for Self-Evolving World Models

Xin Wang, Wenxuan Liu, Tongtong Feng +1

The paper introduces autonomous video generation, a framework that creates future video frames conditioned on counterfactual interventions and embodiment constraints, enabling self…

cs.RO2026

EvolvingAgent: Curriculum Self-evolving Agent with Continual World Model for Long-Horizon Tasks

Tongtong Feng, Xin Wang, Zekai Zhou +5

Completing Long-Horizon (LH) tasks in open-ended worlds is an important yet difficult problem for embodied agents. Existing approaches suffer from two key challenges: (1) they heav…

cs.RO2026

ALAS: Adaptive Long-Horizon Action Synthesis via Async-pathway Stream Disentanglement

Yutong Shen, Hangxu Liu, Lei Zhang +4

Long-Horizon (LH) tasks in Human-Scene Interaction (HSI) are complex multi-step tasks that require continuous planning, sequential decision-making, and extended execution across do…

cs.ET2026

Self-evolving Embodied AI

Tongtong Feng, Xin Wang, Wenwu Zhu

Embodied Artificial Intelligence (AI) is an intelligent system formed by agents and their environment through active perception, embodied cognition, and action interaction. Existin…

cs.CV2025

PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement

Yu-Wei Zhan, Xin Wang, Hong Chen +6

Video Large Language Models (Video LLMs) have shown impressive performance across a wide range of video-language tasks. However, they often fail in scenarios requiring a deeper und…

cs.AI2025

BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models

Yu-Wei Zhan, Xin Wang, Pengzhe Mao +3

Building generalist embodied agents requires a unified system that can interpret multimodal goals, model environment dynamics, and execute reliable actions across diverse real-worl…