works on

From the 1 of 14 linked papers with an AI index.

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Haodong Yan, Junfeng Li, Junjie He +12

Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs…

cs.CV2026

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Zehua Fan, Junjie He, Wenxuan Song +14

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…

cs.CV2026

Video = World + Event Stream

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +24

The paper introduces Wan-Streamer v0.3, a model that treats video as a combination of a persistent world and a dynamic event stream, enabling real-time multimodal audio‑visual inte…

cs.CV2026

Wan-Streamer v0.2: Higher Resolution, Same Latency

Lianghua Huang, Zhi-Fan Wu, Yupeng Shi +23

We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises…

cs.CV2026

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Lianghua Huang, Zhi-Fan Wu, Wei Wang +22

We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex audio-visual interaction. W…

cs.CV2026

DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models

Zhide Zhong, Junfeng Li, Junjie He +10

Vision-Language-Action (VLA) models map visual observations and language instructions directly to robotic actions. While effective for simple tasks, standard VLA models often strug…