5 papers
OccFace: Unified Occlusion-Aware Facial Landmark Detection with Per-Point Visibility
Xinhao Xiang, Zhengxin Li, Saurav Dhakad +3
Accurate facial landmark detection under occlusion remains challenging, especially for human-like faces with large appearance variation and rotation-driven self-occlusion. Existing…
EEPNet-V2: Patch-to-Pixel Solution for Efficient Cross-Modal Registration between LiDAR Point Cloud and Camera Image
Yuanchao Yue, Hui Yuan, Zhengxin Li +2
The primary requirement for cross-modal data fusion is the precise alignment of data from different sensors. However, the calibration between LiDAR point clouds and camera images i…
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
Parker Liu, Chenxin Li, Zhengxin Li +7
Vision-language models (VLMs) excel at descriptive tasks, but whether they truly understand scenes from visual observations remains uncertain. We introduce IR3D-Bench, a benchmark…
Feature Compression for Cloud-Edge Multimodal 3D Object Detection
Chongzhen Tian, Zhengxin Li, Hui Yuan +3
Machine vision systems, which can efficiently manage extensive visual perception tasks, are becoming increasingly popular in industrial production and daily life. Due to the challe…
The Third Monocular Depth Estimation Challenge
Jaime Spencer, Fabio Tosi, Matteo Poggi +38
This paper discusses the results of the third edition of the Monocular Depth Estimation Challenge (MDEC). The challenge focuses on zero-shot generalization to the challenging SYNS-…