6 papers
RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection
Zhihao Zhang, Gengwei Zhang, Tianlong Chen +1
Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories…
Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
Zhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan +1
Monocular 3D detection (Mono3D) aims to infer 3D bounding boxes from a single RGB image. Without auxiliary sensors such as LiDAR, this task is inherently ill-posed since the 3D-to-…
Towards Intrinsic-Aware Monocular 3D Object Detection
Zhihao Zhang, Abhinav Kumar, Xiaoming Liu
Monocular 3D object detection (Mono3D) aims to infer object locations and dimensions in 3D space from a single RGB image. Despite recent progress, existing methods remain highly se…
EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
Huaizhi Qu, Ruichen Zhang, Shuqing Luo +5
Recent advances in foundation models have driven remarkable progress in image editing, yet their extension to 3D editing remains underexplored. A natural approach is to replace the…
Advancing Real-World Parking Slot Detection with Large-Scale Dataset and Semi-Supervised Baseline
Zhihao Zhang, Chunyu Lin, Lang Nie +2
As automatic parking systems evolve, the accurate detection of parking slots has become increasingly critical. This study focuses on parking slot detection using surround-view came…
CHARM3R: Towards Unseen Camera Height Robust Monocular 3D Detector
Abhinav Kumar, Yuliang Guo, Zhihao Zhang +3
Monocular 3D object detectors, while effective on data from one ego camera height, struggle with unseen or out-of-distribution camera heights. Existing methods often rely on Plucke…