activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Physical Object Understanding with a Physically Controllable World Model

Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen +9

A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solvi…

cs.CV2026

Unified 3D Scene Understanding Through Physical World Modeling

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh +4

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essent…

cs.CV2026

Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

Yutong Zhang, Jiaxin Chen, Honglin Chen +5

Memory-efficient transfer learning (METL) approaches have recently achieved promising performance in adapting pre-trained models to downstream tasks. They avoid applying gradient b…

cs.CV2025

World Modeling with Probabilistic Structure Integration

Klemen Kotar, Wanhee Lee, Rahul Venkatesh +13

We present Probabilistic Structure Integration (PSI), a system for learning richly controllable and flexibly promptable world models from data. PSI consists of a three-step cycle.…

cs.CV2025

Discovering and using Spelke segments

Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen +10

Segments in computer vision are often defined by semantic considerations and are highly dependent on category-specific conventions. In contrast, developmental psychology suggests t…

cs.CV2025

3D Scene Understanding Through Local Random Access Sequence Modeling

Wanhee Lee, Klemen Kotar, Rahul Mysore Venkatesh +4

3D scene understanding from single images is a pivotal problem in computer vision with numerous downstream applications in graphics, augmented reality, and robotics. While diffusio…