8 papers
Semantic Image Synthesis via Diffusion Models
Wengang Zhou, Weilun Wang, Jianmin Bao +4
Denoising Diffusion Probabilistic Models (DDPMs) have achieved remarkable success in various image generation tasks compared with Generative Adversarial Nets (GANs). Recent work on…
Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
Daichi Zhang, Tong Zhang, Jianmin Bao +2
With the rapid development of generative models, detecting generated fake images to prevent their malicious use has become a critical issue recently. Existing methods frame this ch…
HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
Wangzheng Shi, Yinglin Zheng, Yuxin Lin +3
Hair transfer is increasingly valuable across domains such as social media, gaming, advertising, and entertainment. While significant progress has been made in single-image hair tr…
SmartEraser: Remove Anything from Images using Masked-Region Guidance
Longtao Jiang, Zhendong Wang, Jianmin Bao +5
Object removal has so far been dominated by the mask-and-inpaint paradigm, where the masked region is excluded from the input, leaving models relying on unmasked areas to inpaint t…
Fast Autoregressive Models for Continuous Latent Generation
Tiankai Hang, Jianmin Bao, Fangyun Wei +1
Autoregressive models have demonstrated remarkable success in sequential data generation, particularly in NLP, but their extension to continuous-domain image generation presents si…
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Microsoft, :, Abdelrahman Abouelenin +73
We introduce Phi-4-Mini and Phi-4-Multimodal, compact yet highly capable language and multimodal models. Phi-4-Mini is a 3.8-billion-parameter language model trained on high-qualit…