works on

From the 1 of 54 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

44 papers · 1 filter

cs.CV2026

OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection

Mariia Gladkova, Neehar Peri, Ishan Khatri +2

Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at…

cs.CV2026

LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

Zixin Guo, Yehonathan Litman, Yifeng He +3

LightCrafter introduces a hybrid method that first renders a video with physically‑based rendering under the target lighting and then refines it with a diffusion model, enabling co…

cs.CV2026

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

Gautam Rajendrakumar Gare, Neehar Peri, Matvei Popov +3

Multi-Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW-13 and RefCOCO. However, state-of-the-art models still s…

cs.CV2026

Steerable Visual Representations

Jona Ruthardt, Manu Gaur, Deva Ramanan +2

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification,…

cs.CV2026

UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Vehicles

Siyi Li, Qingwen Zhang, Ishan Khatri +4

LiDAR scene flow is the task of estimating per-point 3D motion between consecutive point clouds. Recent methods achieve centimeter-level accuracy on popular autonomous vehicle (AV)…

cs.CV2026

Modality Forcing for Scalable Spatial Generation

Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski +2

Text-to-image (T2I) models contain rich spatial priors. Synthesizing photorealistic, cluttered scenes requires an understanding of geometry, including perspective and relative scal…