activity
20242026
collaborators

12 papers

cs.RO2026

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

Yutong Hu, Fengjiao Chen, Xuezhi Cao +1

Long-horizon robot manipulation requires memory, but not necessarily inside the action policy. To address such tasks, current agentic systems often combine VLAs with planners and g…

cs.RO2026

FOCI Policy: Focus on Object-Centric Interactions for Relational Manipulation Policies

Ze Fu, Pinhao Song, Yutong Hu +1

Object-centric manipulation policies improve generalization by modeling object motion instead of directly predicting robot actions. However, existing methods are often limited by r…

cs.RO2026

Assistron: Bayesian Shared Autonomy with Off-the-shelf Vision-Language-Action Models

Pinhao Song, Ze Fu, Yutong Hu +1

We propose Assistron, a shared autonomy model that leverages Vision-Language-Action (VLA) models to assist the user in daily activities. Our approach is grounded in two core princi…

cs.CV2026

MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization

Yutong Hu, Siyuan Tan, Shaocheng Yan +3

Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both intuitive and structured. Althou…

cs.AI2026

FF-JEPA: Long-Horizon Planning in World Models with Latent Planners

Sergi Masip, Jonathan Swinnen, Yutong Hu +2

Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods…

cs.LG2026

ELVIS: Ensemble-Calibrated Latent Imagination for Long-Horizon Visual MPC

Yurui Du, Pinhao Song, Yutong Hu +1

A central challenge of visual control with model-based reinforcement learning (RL) is reliable long-horizon planning: long rollouts with learned latent dynamics exhibit branching f…