223 citations · 388 across the 16 of their papers we have counts for
5 papers · 1 filter
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Mariam Hassan, Sebastian Stapf, Ahmad Rahimi +17
We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, ou…
Reinforcement Learning Meets Visual Odometry
Nico Messikommer, Giovanni Cioffi, Mathias Gehrig +1
Visual Odometry (VO) is essential to downstream mobile robotics and augmented/virtual reality tasks. Despite recent advances, existing VO methods still rely on heuristic design cho…
State Space Models for Event Cameras
Nikola Zubić, Mathias Gehrig, Davide Scaramuzza
Today, state-of-the-art deep neural networks that process event-camera data first convert a temporal window of events into dense, grid-like input representations. As such, they exh…
Seeing Behind Dynamic Occlusions with Event Cameras
Rong Zou, Manasi Muglikar, Nico Messikommer +1
Unwanted camera occlusions, such as debris, dust, rain-drops, and snow, can severely degrade the performance of computer-vision systems. Dynamic occlusions are particularly challen…
COVERED, CollabOratiVE Robot Environment Dataset for 3D Semantic segmentation
Charith Munasinghe, Fatemeh Mohammadi Amin, Davide Scaramuzza +1
Safe human-robot collaboration (HRC) has recently gained a lot of interest with the emerging Industry 5.0 paradigm. Conventional robots are being replaced with more intelligent and…