14 citations · 14 across the 3 of their papers we have counts for
Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning
Yi Li, Alexandre Chapin, Liming Chen +2
Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features…
cs.RO2026
Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
Yudong Liu, Yuan Li, Zijia Tang +12
Dual-system Vision-Language-Action (VLA) models achieve state-of-the-art robotic manipulation but are bottlenecked by the VLM backbone, which must execute at every control step whi…