9 papers
Geometry-Editable and Appearance-Preserving Object Compositon
Jianman Lin, Haojie Li, Chunmei Qing +3
General object composition (GOC) aims to seamlessly integrate a target object into a background scene with desired geometric properties, while simultaneously preserving its fine-gr…
Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Yujie Zhu, Jianman Lin +4
Speech-preserving facial expression manipulation (SPFEM) aims to enhance human expressiveness without altering mouth movements tied to the original speech. A primary challenge in t…
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
Zijian Song, Qichang Li, Sihan Qin +4
The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning…
In-Situ Tweedie Discrete Diffusion Models
Xiao Li, Jiaqi Zhang, Shuxiang Zhang +3
While diffusion models excel at generating continuous data such as images, adapting them to discrete tasks has relied on indirect approaches that either operate in continuous embed…
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
Zijian Song, Sihan Qin, Tianshui Chen +2
The scarcity of manipulation data has motivated the use of pretrained large models from other modalities in robotics. In this work, we build upon autoregressive video generation mo…
Neural Scene Designer: Self-Styled Semantic Image Manipulation
Jianman Lin, Tianshui Chen, Chunmei Qing +4
Maintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing…