collaborators

6 papers

eess.IV2025

Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large Motion

Junru Lin, Chirag Vashist, Mikaela Angelina Uy +4

Existing dynamic scene interpolation methods typically assume that the motion between consecutive timesteps is small enough so that displacements can be locally approximated by lin…

cs.GR2025

Mixture of Contexts for Long Video Generation

Shengqu Cai, Ceyuan Yang, Lvmin Zhang +10

Long video generation is fundamentally a long context memory problem: models must retain and retrieve salient events across a long range without collapsing or drifting. However, sc…

cs.CV2025

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

Phillip Y. Lee, Jihyeon Je, Chanho Park +3

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environmen…

cs.GR2025

GenAnalysis: Joint Shape Analysis by Learning Man-Made Shape Generators with Deformation Regularizations

Yuezhi Yang, Haitao Yang, Kiyohiro Nakayama +3

We present GenAnalysis, an implicit shape generation framework that allows joint analysis of man-made shapes, including shape matching and joint shape segmentation. The key idea is…

cs.CV2024

Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules

Yiqing Liang, Mikhail Okunev, Mikaela Angelina Uy +4

Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is int…

cs.CV2024

ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling

William Yicheng Zhu, Keren Ye, Junjie Ke +4

Recognizing and disentangling visual attributes from objects is a foundation to many computer vision applications. While large vision language representations like CLIP had largely…