5 papers
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
Shizhe Chen, Paul Pacaud, Cordelia Schmid
Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most exi…
BrickNet: Graph-Backed Generative Brick Assembly
Peter Kulits, Cordelia Schmid
We train a language model to generate LEGO-brick build sequences. While prior work has been restricted to discrete, voxel-like towers, we consider a much broader set of pieces, enc…
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +3
Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plau…
MetricNet: Recovering Metric Scale in Generative Navigation Policies
Abhijeet Nayak, Débora Oliveira Makowski, Samiran Gode +2
Generative navigation policies have made rapid progress in improving end-to-end learned navigation. Despite their promising results, this paradigm has two structural problems. Firs…
FlowNav: Combining Flow Matching and Depth Priors for Efficient Navigation
Samiran Gode, Abhijeet Nayak, Débora N. P. Oliveira +3
Effective robot navigation in unseen environments is a challenging task that requires precise control actions at high frequencies. Recent advances have framed it as an image-goal-c…