From the 1 of 65 linked papers with an AI index.
3 citations · 4 across the 20 of their papers we have counts for
19 papers · 1 filter
Visual General Intelligence: A White Paper
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…
OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection
Mariia Gladkova, Neehar Peri, Ishan Khatri +2
Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at…
LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting
Zixin Guo, Yehonathan Litman, Yifeng He +3
LightCrafter introduces a hybrid method that first renders a video with physically‑based rendering under the target lighting and then refines it with a diffusion model, enabling co…
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
Gautam Rajendrakumar Gare, Neehar Peri, Matvei Popov +3
Multi-Modal LLMs (MLLMs) demonstrate strong visual grounding capabilities on popular object detection benchmarks like OdinW-13 and RefCOCO. However, state-of-the-art models still s…
Steerable Visual Representations
Jona Ruthardt, Manu Gaur, Deva Ramanan +2
Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification,…