activity
20242026
most citedUrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

17 citations · 22 across the 14 of their papers we have counts for

collaborators

14 papers

cs.AI2026

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

Rongze Tang, Jianjie Fang, Zhaolu Wang +8

World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing appr…

cs.AI2026

CAER: Causal Action Effect Reweighting for World Model Training

Jianjie Fang, Xvyuan Liu, Ziyou Wang +9

World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agen…

cs.CV2026

DensityKV: Density-Guided KV Cache Compression for Long Video Generation

Wenqu Zhao, Xuemin Chi, Xin Zhang +6

Autoregressive video diffusion models enable streaming generation through sliding-window attention, but each generated block is conditioned on previously generated content, causing…

cs.RO2026

WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory and In-Context Learning

Haisheng Su, Zongdai Liu, Xin Jin +13

World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constraine…

cs.RO2026

Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control

Jianjie Fang, Yongyan Xu, Ziyou Wang +13

World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, fo…

cs.RO2026

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

Baining Zhao, Jiacheng Xu, Weicheng Feng +13

Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial V…