collaborators

7 papers

cs.CV2026

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

Ananya Bal, Kartik Sharma, Ethan Lai +4

Methods for text-based generation of hand-object interaction (HOI) sequences primarily focus on producing smooth, physically plausible trajectories. A truly utilitarian method shou…

cs.CV2026

JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation

Mingyeong Song, Jungbin Cho, Jisoo Kim +5

Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challengi…

cs.AI2026

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

Chunjiang Liu, Xiaoyuan Wang, Haoyu Chen +3

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, i…

cs.CV2026

GeoStream: Toward Precise Camera Controlled Streaming Video Generation

Yizhou Zhao, Yifan Wang, Xiaoyuan Wang +11

Accurate interactive camera control is essential for video-based world models, but most existing approaches learn camera motion implicitly, leading to inaccurate control under out-…

cs.CV2026

Accelerating Vision Transformers with Adaptive Patch Sizes

Rohan Choudhury, JungEun Kim, Jinhyung Park +3

Vision Transformers (ViTs) partition input images into uniformly sized patches regardless of their content, resulting in long input sequence lengths for high-resolution images. We…

cs.CV2026

MOSIV: Multi-Object System Identification from Videos

Chunjiang Liu, Xiaoyuan Wang, Qingran Lin +9

We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discret…