papers

Publications (17)

cs.RO2023

PolyMerge: A Novel Technique aimed at Dynamic HD Map Updates Leveraging Polylines

Mohamed Sayed, Stepan Perminov, Dzmitry Tsetserukou

Currently, High-Definition (HD) maps are a prerequisite for the stable operation of autonomous vehicles. Such maps contain information about all static road objects for the vehicle…

cs.CV2024

GroundUp: Rapid Sketch-Based 3D City Massing

Gizem Esra Unlu, Mohamed Sayed, Yulia Gryaditskaya +1

We propose GroundUp, the first sketch-based ideation tool for 3D city massing of urban areas. We focus on early-stage urban design, where sketching is a common tool and the design…

cs.CV2026

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

Matias Turkulainen, Akshay Krishnan, Filippo Aleotti +6

We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level AND by satellite. Faithful reconstruct…

cs.GR2022

LookOut! Interactive Camera Gimbal Controller for Filming Long Takes

Mohamed Sayed, Robert Cinca, Enrico Costanza +1

The job of a camera operator is challenging, and potentially dangerous, when filming long moving camera shots. Broadly, the operator must keep the actors in-frame while safely navi…

cs.CV2024

AirPlanes: Accurate Plane Estimation via 3D-Consistent Embeddings

Jamie Watson, Filippo Aleotti, Mohamed Sayed +5

Extracting planes from a 3D scene is useful for downstream tasks in robotics and augmented reality. In this paper we tackle the problem of estimating the planar surfaces in a scene…

cs.CV2025

MVSAnywhere: Zero-Shot Multi-View Stereo

Sergio Izquierdo, Mohamed Sayed, Michael Firman +6

Computing accurate depth from multiple views is a fundamental and longstanding challenge in computer vision. However, most existing approaches do not generalize well across differe…

cs.RO2026

Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning

Mohamed Sayed, Wolfram Burgard, Tanja Katharina Kaiser

Cooperative object transportation is essential in numerous domains, including industrial to domestic services. A popular transportation strategy is to carry objects on top of multi…

cs.CV2025

Complete Gaussian Splats from a Single Image with Denoising Diffusion Models

Ziwei Liao, Mohamed Sayed, Steven L. Waslander +3

Gaussian splatting typically requires dense observations of the scene and can fail to reconstruct occluded and unobserved areas. We propose a latent diffusion model to reconstruct…

cs.CV2023

Virtual Occlusions Through Implicit Depth

Jamie Watson, Mohamed Sayed, Zawar Qureshi +4

For augmented reality (AR), it is important that virtual assets appear to `sit among' real world objects. The virtual element should variously occlude and be occluded by real matte…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.CV2021

Improved Handling of Motion Blur in Online Object Detection

Mohamed Sayed, Gabriel Brostow

We wish to detect specific categories of objects, for online vision systems that will run in the real world. Object detection is already very challenging. It is even harder when th…

cs.CV2022

Interactive Sketching of Mannequin Poses

Gizem Unlu, Mohamed Sayed, Gabriel Brostow

It can be easy and even fun to sketch humans in different poses. In contrast, creating those same poses on a 3D graphics "mannequin" is comparatively tedious. Yet 3D body poses are…

cs.CV2022

SimpleRecon: 3D Reconstruction Without 3D Convolutions

Mohamed Sayed, John Gibson, Jamie Watson +3

Traditionally, 3D indoor scene reconstruction from posed images happens in two phases: per-image depth estimation, followed by depth merging and surface reconstruction. Recently, a…

cs.CV2024

An Empirical Study of the Generalization Ability of Lidar 3D Object Detectors to Unseen Domains

George Eskandar, Chongzhe Zhang, Abhishek Kaushik +3

3D Object Detectors (3D-OD) are crucial for understanding the environment in many robotic tasks, especially autonomous driving. Including 3D information via Lidar sensors improves…

cs.RO2024

MoveTouch: Robotic Motion Capturing System with Wearable Tactile Display to Achieve Safe HRI

Ali Alabbas, Miguel Altamirano Cabrera, Mohamed Sayed +3

The collaborative robot market is flourishing as there is a trend towards simplification, modularity, and increased flexibility on the production line. But when humans and robots a…

cs.CV2024

DoubleTake: Geometry Guided Depth Estimation

Mohamed Sayed, Filippo Aleotti, Jamie Watson +5

Estimating depth from a sequence of posed RGB images is a fundamental computer vision task, with applications in augmented reality, path planning etc. Prior work typically makes us…

cs.CV2025

Morpheus: Text-Driven 3D Gaussian Splat Shape and Color Stylization

Jamie Wynn, Zawar Qureshi, Jakub Powierza +2

Exploring real-world spaces using novel-view synthesis is fun, and reimagining those worlds in a different style adds another layer of excitement. Stylized worlds can also be used…