5 papers
Per-Group Error, Not Total MSE: Fine-Tuning Vision-Language-Action Models for 11-DoF Mobile Manipulation
Pau Montagut Bofi, Mario GarcÃa Blasco, Tessa Pulli +1
Fine-tuning Vision-Language-Action (VLA) models for mobile manipulators with heterogeneous joint spaces can produce a counterintuitive result: the checkpoint with the lowest aggreg…
Hierarchical Prompting with Dual LLM Modules for Robotic Task and Motion Planning
Karolina Źróbek, Tessa Pulli, PaweŠGajewski +2
We present a hierarchical language-driven framework for robotic task and motion planning to improve natural, intuitive human-robot interaction in service and assistance scenarios.…
OSCAR: Open-Set CAD Retrieval from a Language Prompt and a Single Image
Tessa Pulli, Jean-Baptiste Weibel, Peter Hönig +3
6D object pose estimation plays a crucial role in scene understanding for applications such as robotics and augmented reality. To support the needs of ever-changing object sets in…
Multi-Modal 3D Mesh Reconstruction from Images and Text
Melvin Reka, Tessa Pulli, Markus Vincze
6D object pose estimation for unseen objects is essential in robotics but traditionally relies on trained models that require large datasets, high computational costs, and struggle…
Enhancing Transparent Object Pose Estimation: A Fusion of GDR-Net and Edge Detection
Tessa Pulli, Peter Hönig, Stefan Thalhammer +2
Object pose estimation of transparent objects remains a challenging task in the field of robot vision due to the immense influence of lighting, background, and reflections. However…