activity
20242026
collaborators

13 papers

cs.CV2026

OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection

Mariia Gladkova, Neehar Peri, Ishan Khatri +2

Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at…

cs.CV2026

TRASE: Tracking-free 4D Segmentation and Editing

Yun-Jin Li, Mariia Gladkova, Yan Xia +1

Understanding dynamic 3D scenes is crucial for extended reality (XR) and autonomous driving. Incorporating semantic information into 3D reconstruction enables holistic scene repres…

cs.RO2025

ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation

Yuhan Wu, Tiantian Wei, Shuo Wang +4

Interactive articulated manipulation requires long-horizon, multi-step interactions with appliances while maintaining physical consistency. Existing vision-language and diffusion-b…

cs.CV2025

Text2Loc++: Generalizing 3D Point Cloud Localization from Natural Language

Yan Xia, Letian Shi, Yilin Di +2

We tackle the problem of localizing 3D point cloud submaps using complex and diverse natural language descriptions, and present Text2Loc++, a novel neural network designed for effe…

cs.CV2025

OPAL: Visibility-aware LiDAR-to-OpenStreetMap Place Recognition via Adaptive Radial Fusion

Shuhao Kang, Martin Y. Liao, Yan Xia +3

LiDAR place recognition is a critical capability for autonomous navigation and cross-modal localization in large-scale outdoor environments. Existing approaches predominantly depen…

cs.CV2025

True Multimodal In-Context Learning Needs Attention to the Visual Context

Shuo Chen, Jianzhe Liu, Zhen Han +5

Multimodal Large Language Models (MLLMs), built on powerful language backbones, have enabled Multimodal In-Context Learning (MICL)-adapting to new tasks from a few multimodal demon…