7 papers
Training-free Motion Factorization for Compositional Video Generation
Zixuan Wang, Ziqin Zhou, Feng Chen +4
Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting t…
Enhancing Text-to-Image Generation via End-Edge Collaborative Hybrid Super-Resolution
Chongbin Yi, Yuxin Liang, Ziqi Zhou +1
Artificial Intelligence-Generated Content (AIGC) has made significant strides, with high-resolution text-to-image (T2I) generation becoming increasingly critical for improving user…
Visual Self-Refinement for Autoregressive Models
Jiamian Wang, Ziqi Zhou, Chaithanya Kumar Mummadi +5
Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequen…
Cutting the Skip: Training Residual-Free Transformers
Yiping Ji, James Martens, Jianqiao Zheng +5
Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without skip (residual) connectio…
TATTOO: Training-free AesTheTic-aware Outfit recOmmendation
Yuntian Wu, Xiaonan Hu, Ziqi Zhou +1
The global fashion e-commerce market relies significantly on intelligent and aesthetic-aware outfit-completion tools to promote sales. While previous studies have approached the pr…
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
Zicheng Duan, Yuxuan Ding, Chenhui Gou +3
Zero-shot personalized image generation models aim to produce images that align with both a given text prompt and subject image, requiring the model to incorporate both sources of…