activity
20242026
collaborators

9 papers

cs.AI2026

Foresight Without Seeing: Latent Futures for World Action Models

Jiakai Huang, Zhongbo Wu, Zheng Zhang +3

World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction. Existing WAMs…

cs.AI2026

Reward as An Agent for Embodied World Models

Pu Li, Zhigang Lin, Qiang Wu +3

While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training distribution, limiting exploration, behavio…

cs.AI2026

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

Kairos Team, Fei Wang, Shan You +21

We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully si…

cs.CV2026

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

Chenyu Hui, Xiaodi Huang, Siyu Xu +5

Vision-language-action (VLA) models typically rely on large-scale real-world videos, whereas simulated data, despite being inexpensive and highly parallelizable to collect, often s…

cs.CV2025

Learning to Expand Images for Efficient Visual Autoregressive Modeling

Ruiqing Yang, Kaixin Zhang, Zheng Zhang +2

Autoregressive models have recently shown great promise in visual generation by leveraging discrete token sequences akin to language modeling. However, existing approaches often su…

cs.CV2025

ActVAR: Activating Mixtures of Weights and Tokens for Efficient Visual Autoregressive Generation

Kaixin Zhang, Ruiqing Yang, Yuan Zhang +2

Visual Autoregressive (VAR) models enable efficient image generation via next-scale prediction but face escalating computational costs as sequence length grows. Existing static pru…