29 citations · 67 across the 32 of their papers we have counts for
35 papers
PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding
Siru Zhong, Qiongyan Wang, Xiaohui Lv +5
Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value…
WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
Peterson Co, Sicheng Hu, Chunxuan Jiao +17
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this prom…
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
Xiangkai Ma, Yue Ma, Junjie Wang +6
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visua…
ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching
Zihan Liu, Yuzhe Zhuang, Yuanzu Li +4
JEPA-style visual world models offer an effective paradigm for visual goal planning by predicting future latent representations. Existing methods typically learn local transition c…
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
Xinyan Cai, Shiguang Wu, Dafeng Chi +4
In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imag…
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
Min Lin, Xiwen Liang, Bingqian Lin +12
Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly…