activity
20192026
most citedA Survey on Vision-Language-Action Models for Embodied AI

29 citations · 67 across the 32 of their papers we have counts for

collaborators

35 papers

cs.CV2026

PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

Siru Zhong, Qiongyan Wang, Xiaohui Lv +5

Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value…

cs.RO2026

WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation

Peterson Co, Sicheng Hu, Chunxuan Jiao +17

Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this prom…

cs.RO2026

Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models

Xiangkai Ma, Yue Ma, Junjie Wang +6

World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visua…

cs.AI2026

ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching

Zihan Liu, Yuzhe Zhuang, Yuanzu Li +4

JEPA-style visual world models offer an effective paradigm for visual goal planning by predicting future latent representations. Existing methods typically learn local transition c…

cs.CV2025

EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

Xinyan Cai, Shiguang Wu, Dafeng Chi +4

In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imag…

cs.RO2025

EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation

Min Lin, Xiwen Liang, Bingqian Lin +12

Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly…