23 citations · 145 across the 67 of their papers we have counts for
9 papers · 1 filter
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
Yuyou Zhang, Radu Corcodel, Chiori Hori +2
We present SpinBench, a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision language models (VLMs). SpinBench is designed around the core challenge…
Feature-EndoGaussian: Feature Distilled Gaussian Splatting in Surgical Deformable Scene Reconstruction
Kai Li, Junhao Wang, William Han +1
Minimally invasive surgery (MIS) requires high-fidelity, real-time visual feedback of dynamic and low-texture surgical scenes. To address these requirements, we introduce FeatureEn…
MONA: Moving Object Detection from Videos Shot by Dynamic Camera
Boxun Hu, Mingze Xia, Ding Zhao +1
Dynamic urban environments, characterized by moving cameras and objects, pose significant challenges for camera trajectory estimation by complicating the distinction between camera…
The RoboDepth Challenge: Methods and Advancements Towards Robust Depth Estimation
Lingdong Kong, Yaru Niu, Shaoyuan Xie +39
Accurate depth estimation under out-of-distribution (OoD) scenarios, such as adverse weather conditions, sensor failure, and noise contamination, is desirable for safety-critical a…
MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos
Jielin Qiu, Jiacheng Zhu, William Han +9
Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets,…
Robustness Certification of Visual Perception Models via Camera Motion Smoothing
Hanjiang Hu, Zuxin Liu, Linyi Li +2
A vast literature shows that the learning-based visual perception model is sensitive to adversarial noises, but few works consider the robustness of robotic perception models under…