8 papers
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models
Yujun Tong, Dongliang Chang, Zijin Yin +3
The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Despite recent works such as BA…
Generative Visual Chain-of-Thought for Image Editing
Zijin Yin, Tiankai Hang, Yiji Cheng +9
Existing image editing methods struggle to perceive where to edit, especially under complex scenes and nuanced spatial instructions. To address this issue, we propose Generative Vi…
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
Zijin Yin, Bing Li, Kongming Liang +4
Semantic segmentation takes pivotal roles in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is criti…
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
Jiayi Gao, Zijin Yin, Changcheng Hua +5
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have t…
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
Runpu Wei, Zijin Yin, Shuo Zhang +8
Inpainting algorithms have achieved remarkable progress in removing objects from images, yet still face two challenges: 1) struggle to handle the object's visual effects such as sh…
PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Image Segmentation
Zhonghao Yan, Zijin Yin, Tianyu Lin +3
The Segment Anything Model (SAM) has demonstrated strong and versatile segmentation capabilities, along with intuitive prompt-based interactions. However, customizing SAM for medic…