4 papers
DINOcular: Self-Supervised Visuospatial Representations
Farkhat Almukhamedov, Sami Azirar, Hermann Blum
We introduce a self-supervised framework for learning joint visuospatial representations from RGB-D observations. While modern vision foundation models are trained almost exclusive…
ContactFlow: A video action conditioning that transfers across embodiments
Sami Azirar, Enrico Pallotta, Jan Nogga +3
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world…
SYMBOLIZER: Symbolic Model-free Task Planning with VLMs
Sami Azirar, Zlatan Ajanovic, Hermann Blum
Traditional Task and Motion Planning (TAMP) systems depend on physics models for motion planning and discrete symbolic models for task planning. Although physics model are often av…
IQLS: Framework for leveraging Metadata to enable Large Language Model based queries to complex, versatile Data
Sami Azirar, Hossam A. Gabbar, Chaouki Regoui
As the amount and complexity of data grows, retrieving it has become a more difficult task that requires greater knowledge and resources. This is especially true for the logistics…