collaborators

13 papers

cs.CV2026

OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection

Mariia Gladkova, Neehar Peri, Ishan Khatri +2

Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at…

cs.CV2026

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

Gautam Rajendrakumar Gare, Neehar Peri, Matvei Popov +3

Multi-Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW-13 and RefCOCO. However, state-of-the-art models still s…

cs.RO2026

Bridging Handheld and Teleoperated Supervision for Contact-Rich Manipulation via State-Gated Experts

Vidullan Surendran, Neehar Peri, David Watkins

Handheld data collection systems, such as the Universal Manipulation Interface (UMI), enable scalable data collection across diverse environments but only capture observed actions…

cs.CV2026

UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles

Siyi Li, Qingwen Zhang, Ishan Khatri +4

LiDAR scene flow is the task of estimating per-point 3D motion between consecutive point clouds. Recent methods achieve centimeter-level accuracy on popular autonomous vehicle (AV)…

cs.CV2026

MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion

Zihan Wang, Jeff Tan, Tarasha Khurana +2

We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panopt…

cs.CV2026

RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

Isaac Robinson, Peter Robicheaux, Matvei Popov +2

Open-vocabulary detectors achieve impressive performance on COCO, but often fail to generalize to real-world datasets with out-of-distribution classes not typically found in their…