activity
20242026
collaborators

8 papers

cs.CV2026

Unified 3D Scene Understanding Through Physical World Modeling

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh +4

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essent…

cs.CV2026

Characterizing the visual representation of objects from the child's view

Jane Yang, Tarun Sepuri, Alvin Wei Ming Tan +3

Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning process look like? We analyzed…

cs.AI2026

Zero-shot World Models Are Developmentally Efficient Learners

Khai Loong Aw, Klemen Kotar, Wanhee Lee +6

Young children demonstrate early abilities to understand their physical world, estimating depth, motion, object coherence, interactions, and many other aspects of physical scene un…

cs.CV2025

Taming generative video models for zero-shot optical flow extraction

Seungwoo Kim, Khai Loong Aw, Klemen Kotar +8

Extracting optical flow from videos remains a core computer vision problem. Motivated by the recent success of large general-purpose models, we ask whether frozen self-supervised v…

cs.CV2025

Assessing the alignment between infants' visual and linguistic experience using multimodal language models

Alvin Wei Ming Tan, Jane Yang, Tarun Sepuri +6

Figuring out which objects or concepts words refer to is a central language learning challenge for young children. Most models of this process posit that children learn early objec…

cs.CV2025

World Modeling with Probabilistic Structure Integration

Klemen Kotar, Wanhee Lee, Rahul Venkatesh +13

We present Probabilistic Structure Integration (PSI), a system for learning richly controllable and flexibly promptable world models from data. PSI consists of a three-step cycle.…