9 papers · 1 filter
PRISM: Progressive Reasoning through Iterative Slot Memory for Vision
Ziyu Wang, Shuangpeng Han, Mengmi Zhang
Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representations under incomplete obser…
Learning to Perceive "Where": Spatial Pretext Tasks for Robust Self-Supervised Learning
Yang Shen, Yusen Cai, Weronika Hryniewska-Guzik +2
Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationships among object parts. To ad…
Learning to See Through a Baby's Eyes: Early Visual Diets Enable Robust Visual Intelligence in Humans and Machines
Yusen Cai, Qing Lin, Bhargava Satya Nunna +1
Newborns perceive the world with low-acuity, color-degraded, and temporally continuous vision, which gradually sharpens as infants develop. To explore the ecological advantages of…
Peering into the Unknown: Active View Selection with Neural Uncertainty Maps for 3D Reconstruction
Zhengquan Zhang, Feng Xu, Mengmi Zhang
Some perspectives naturally provide more information than others. How can an AI system determine which viewpoint offers the most valuable insight for accurate and efficient 3D obje…
Unforgettable Lessons from Forgettable Images: Intra-Class Memorability Matters in Computer Vision
Jie Jing, Yongjian Huang, Serena J. -W. Wang +5
We introduce intra-class memorability, where certain images within the same class are more memorable than others despite shared category characteristics. To investigate what featur…
Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
Cristian Meo, Akihiro Nakano, Mircea Lică +7
Unsupervised object-centric learning from videos is a promising approach towards learning compositional representations that can be applied to various downstream tasks, such as pre…