4 papers
G3Splat: Geometrically Consistent Generalizable Gaussian Splatting
Mehdi Hosseinzadeh, Shin-Fang Chng, Yi Xu +3
3D Gaussians have become a powerful scene representation for real-time splatting and high-quality novel-view synthesis. This has motivated generalizable splatting -- methods that a…
KITE: Keyframe-Indexed Tokenized Evidence for VLM-Based Robot Failure Analysis
Mehdi Hosseinzadeh, King Hang Wong, Feras Dayoub
We present KITE, a training-free, keyframe-anchored, layout-grounded front-end that converts long robot-execution videos into compact, interpretable tokenized evidence for vision-l…
TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals
Stefan Podgorski, Sourav Garg, Mehdi Hosseinzadeh +3
Visual navigation in robotics traditionally relies on globally-consistent 3D maps or learned controllers, which can be computationally expensive and difficult to generalize across…
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
Mehdi Hosseinzadeh, Ian Reid
In the field of autonomous driving and mobile robotics, there has been a significant shift in the methods used to create Bird's Eye View (BEV) representations. This shift is charac…