12 papers
Leveraging Latent Visual Reasoning in Silence
Dongyao Zhu, Zhen Wang, Xi Xiao +7
Latent visual reasoning involves visual evidence more directly in multimodal reasoning by inserting continuous latent tokens before textual generation. However, the necessity of th…
SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability
Amin Karimi Monsefi, Abolfazl Meyarian, Mridul Khurana +6
Animals are described as effectively camouflaged when they blend seamlessly with their surrounding, yet no standardized quantitative measure of this seamlessness exists. We address…
TaxaAdapter: Vision Taxonomy Models are Key to Fine-grained Image Generation over the Tree of Life
Mridul Khurana, Amin Karimi Monsefi, Justin Lee +9
Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the rema…
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
Sooyoung Jeon, Hongjie Tian, Lemeng Wang +7
Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment environments. While the computer visio…
When the City Teaches the Car: Label-Free 3D Perception from Infrastructure
Zhen Xu, Jinsu Yoo, Cristian Bautista +7
Building robust 3D perception for self-driving still relies heavily on large-scale data collection and manual annotation, yet this paradigm becomes impractical as deployment expand…
On the Feasibility and Opportunity of Autoregressive 3D Object Detection
Zanming Huang, Jinsu Yoo, Sooyoung Jeon +6
LiDAR-based 3D object detectors typically rely on proposal heads with hand-crafted components like anchor assignment and non-maximum suppression (NMS), complicating training and li…