activity
20242026
collaborators

5 papers

cs.CV2026

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian +4

Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent…

cs.CV2026

UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception

Mohammad Mahdavian, Gordon Tan, Binbin Xu +3

We present UniScale, a unified, scale-aware multi-view 3D reconstruction framework for robotic applications that flexibly integrates geometric priors through a modular, semanticall…

cs.CV2025

LS-HAR: Language Supervised Human Action Recognition with Salient Fusion, Construction Sites as a Use-Case

Mohammad Mahdavian, Mohammad Loni, Ted Samuelsson +1

Detecting human actions is a crucial task for autonomous robots and vehicles, often requiring the integration of various data modalities for improved accuracy. In this study, we in…

cs.RO2025

STPOTR: Simultaneous Human Trajectory and Pose Prediction Using a Non-Autoregressive Transformer for Robot Following Ahead

Mohammad Mahdavian, Payam Nikdel, Mahdi TaherAhmadi +1

In this paper, we develop a neural network model to predict future human motion from an observed human motion history. We propose a non-autoregressive transformer architecture to l…

cs.CV2024

TileTracker: Tracking Based Vector HD Mapping using Top-Down Road Images

Mohammad Mahdavian, Mo Chen, Yu Zhang

In this paper, we propose a tracking-based HD mapping algorithm for top-down road images, referred to as tile images. While HD maps traditionally rely on perspective camera images,…