14 citations · 14 across the 3 of their papers we have counts for
4 papers
More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning
Yi Li, Alexandre Chapin, Liming Chen +2
Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix task-relevant and task-irrelevant features…
Conditional Multi-Event Temporal Grounding in Long-Form Video
Yuanhao Zou, Arthad Kulkarni, Lucas Tonanez +12
Multimodal large language models have made rapid progress in video temporal grounding, yet real-world applications routinely require localizing every event that satisfies compositi…
Latent Bridge: Feature Delta Prediction for Efficient Dual-System Vision-Language-Action Model Inference
Yudong Liu, Yuan Li, Zijia Tang +12
Dual-system Vision-Language-Action (VLA) models achieve state-of-the-art robotic manipulation but are bottlenecked by the VLM backbone, which must execute at every control step whi…
Morig: Motion-aware rigging of character meshes from point clouds
Zhan Xu, Yang Zhou, Li Yi +1
We present MoRig, a method that automatically rigs character meshes driven by single-view point cloud streams capturing the motion of performing characters. Our method is also able…