collaborators

8 papers

cs.CV2026

Physical Object Understanding with a Physically Controllable World Model

Rahul Venkatesh, Klemen Kotar, Lilian Naing Chen +9

A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solvi…

cs.CV2026

ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

Chia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang +5

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific…

cs.CV2025

Autoregressive Flow Matching for Motion Prediction

Johnathan Xie, Stefan Stojanov, Cristobal Eyzaguirre +2

Motion prediction has been studied in different contexts with models trained on narrow distributions and applied to downstream tasks in human motion prediction and robotics. Simult…

cs.CV2025

Taming generative video models for zero-shot optical flow extraction

Seungwoo Kim, Khai Loong Aw, Klemen Kotar +8

Extracting optical flow from videos remains a core computer vision problem. Motivated by the recent success of large general-purpose models, we ask whether frozen self-supervised v…

cs.CV2025

World Modeling with Probabilistic Structure Integration

Klemen Kotar, Wanhee Lee, Rahul Venkatesh +13

We present Probabilistic Structure Integration (PSI), a system for learning richly controllable and flexibly promptable world models from data. PSI consists of a three-step cycle.…

cs.CV2025

Weakly-Supervised Learning of Dense Functional Correspondences

Stefan Stojanov, Linan Zhao, Yunzhi Zhang +2

Establishing dense correspondences across image pairs is essential for tasks such as shape reconstruction and robot manipulation. In the challenging setting of matching across diff…