7 papers
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
Zisu Li, Hengye Lyu, Jiaxin Shi +4
Modeling and synthesizing complex hand-object interactions remains a significant challenge, even for state-of-the-art physics engines. Conventional simulation-based approaches rely…
Generative Augmented Reality: Paradigms, Technologies, and Future Applications
Chen Liang, Jiawen Zheng, Yufeng Zeng +7
This paper introduces Generative Augmented Reality (GAR) as a next-generation paradigm that reframes augmentation as a process of world re-synthesis rather than world composition b…
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
Wang Lin, Liyu Jia, Wentao Hu +6
Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapol…
Generalized Visual Relation Detection with Diffusion Models
Kaifeng Gao, Siqi Chen, Hanwang Zhang +3
Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance,…
Seeing World Dynamics in a Nutshell
Qiuhong Shen, Xuanyu Yi, Mingbao Lin +3
We consider the problem of efficiently representing casually captured monocular videos in a spatially- and temporally-coherent manner. While existing approaches predominantly rely…
LoRA of Change: Learning to Generate LoRA for the Editing Instruction from A Single Before-After Image Pair
Xue Song, Jiequan Cui, Hanwang Zhang +4
In this paper, we propose the LoRA of Change (LoC) framework for image editing with visual instructions, i.e., before-after image pairs. Compared to the ambiguities, insufficient s…