209 citations · 919 across the 46 of their papers we have counts for
8 papers · 1 filter
Visual General Intelligence: A White Paper
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
Khiem Vuong, Deva Ramanan, Srinivasa Narasimhan
Rendering views using 3D scene representations such as Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), meshes, or even point clouds produces artifacts when input views ar…
OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection
Mariia Gladkova, Neehar Peri, Ishan Khatri +2
Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at…
LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting
Zixin Guo, Yehonathan Litman, Yifeng He +3
Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic sc…
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
Scaling Short-Term Memory of Visuomotor Policies for Long-Horizon Tasks
Rutav Shah, Rajat Kumar Jenamani, Xiaohan Zhang +5
Many robotic tasks require short-term memory, whether it's retrieving an object that's no longer visible or turning off an appliance after a set period. Yet, most visuomotor polici…