5 papers
Control and Realism: Best of Both Worlds in Layout-to-Image without Training
Bonan Li, Yinhan Hu, Songhua Liu +1
Layout-to-Image generation aims to create complex scenes with precise control over the placement and arrangement of subjects. Existing works have demonstrated that pre-trained Text…
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
Bonan li, Zicheng Zhang, Songhua Liu +2
Visual instruction tuning aims to enable large language models to comprehend the visual world, with a pivotal challenge lying in establishing an effective vision-to-language projec…
LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head Synthesis
Tianqi Li, Ruobing Zheng, Bonan Li +4
Despite significant progress in talking head synthesis since the introduction of Neural Radiance Fields (NeRF), visual artifacts and high training costs persist as major obstacles…
StyO: Stylize Your Face in Only One-shot
Bonan Li, Zicheng Zhang, Xuecheng Nie +4
This paper focuses on face stylization with a single artistic target. Existing works for this task often fail to retain the source content while achieving geometry variation. Here,…
DR-BFR: Degradation Representation with Diffusion Models for Blind Face Restoration
Xinmin Qiu, Bonan Li, Zicheng Zhang +2
Blind face restoration (BFR) is fundamentally challenged by the extensive range of degradation types and degrees that impact model generalization. Recent advancements in diffusion…