collaborators

5 papers

cs.RO2026

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

Shizhe Chen, Paul Pacaud, Cordelia Schmid

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-language backbones. However, most exi…

cs.CV2026

BrickNet: Graph-Backed Generative Brick Assembly

Peter Kulits, Cordelia Schmid

We train a language model to generate LEGO-brick build sequences. While prior work has been restricted to discrete, voxel-like towers, we consider a much broader set of pieces, enc…

cs.CV2026

HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +3

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plau…

cs.RO2026

MetricNet: Recovering Metric Scale in Generative Navigation Policies

Abhijeet Nayak, Débora Oliveira Makowski, Samiran Gode +2

Generative navigation policies have made rapid progress in improving end-to-end learned navigation. Despite their promising results, this paradigm has two structural problems. Firs…

cs.RO2025

FlowNav: Combining Flow Matching and Depth Priors for Efficient Navigation

Samiran Gode, Abhijeet Nayak, Débora N. P. Oliveira +3

Effective robot navigation in unseen environments is a challenging task that requires precise control actions at high frequencies. Recent advances have framed it as an image-goal-c…