A Survey on Personalized Content Synthesis with Diffusion Models
arXiv:2405.05538 · doi:10.1007/s11633-025-1563-3
Abstract
Recent advancements in diffusion models have significantly impacted content creation, leading to the emergence of Personalized Content Synthesis (PCS). By utilizing a small set of user-provided examples featuring the same subject, PCS aims to tailor this subject to specific user-defined prompts. Over the past two years, more than 150 methods have been introduced in this area. However, existing surveys primarily focus on text-to-image generation, with few providing up-to-date summaries on PCS. This paper provides a comprehensive survey of PCS, introducing the general frameworks of PCS research, which can be categorized into test-time fine-tuning (TTF) and pre-trained adaptation (PTA) approaches. We analyze the strengths, limitations, and key techniques of these methodologies. Additionally, we explore specialized tasks within the field, such as object, face, and style personalization, while highlighting their unique challenges and innovations. Despite the promising progress, we also discuss ongoing challenges, including overfitting and the trade-off between subject fidelity and text alignment. Through this detailed overview and analysis, we propose future directions to further the development of PCS.
References in corpus (60)
- FaceNet: A Unified Embedding for Face Recognition and Clustering
- Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models in Vision: A Survey
- Gemini: A Family of Highly Capable Multimodal Models
- DreamFusion: Text-to-3D using 2D Diffusion
- Facial Landmark Detection: a Literature Survey
- Survey on reinforcement learning for language processing
- DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
- Break-A-Scene: Extracting Multiple Concepts from a Single Image
- IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
- P+: Extended Textual Conditioning in Text-to-Image Generation
- InstantID: Zero-shot Identity-Preserving Generation in Seconds
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models
- Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA
- Unified Multi-Modal Latent Diffusion for Joint Subject and Text Conditional Image Generation
- StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation
- AvatarBooth: High-Quality and Customizable 3D Human Avatar Generation
- Highly Personalized Text Embedding for Image Manipulation by Stable Diffusion
- Navigating Text-To-Image Customization: From LyCORIS Fine-Tuning to Model Evaluation
- Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach
- ID-Animator: Zero-Shot Identity-Preserving Human Video Generation
- DreamVideo: Composing Your Dream Videos with Customized Subject and Motion
- Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation
- Identity Encoder for Personalized Diffusion
- Emu3: Next-Token Prediction is All You Need
- Animate124: Animating One Image to 4D Dynamic Scene
- CustomNet: Zero-shot Object Customization with Variable-Viewpoints in Text-to-Image Diffusion Models
- Portrait Diffusion: Training-free Face Stylization with Chain-of-Painting
- CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
- Generate Anything Anywhere in Any Scene
- DreamIdentity: Improved Editability for Efficient Face-identity Preserved Image Generation
- DreamTuner: Single Image is Enough for Subject-Driven Generation
- A Closer Look at Parameter-Efficient Tuning in Diffusion Models
- FaceStudio: Put Your Face Everywhere in Seconds
- InstantFamily: Masked Attention for Zero-shot Multi-ID Image Generation
- Text-Conditional Contextualized Avatars For Zero-Shot Personalization
- InstructBooth: Instruction-following Personalized Text-to-Image Generation
- MC: Multi-concept Guidance for Customized Multi-concept Generation
- ViCo: Plug-and-play Visual Condition for Personalized Text-to-image Generation
- ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning
- Cross Initialization for Personalized Text-to-Image Generation
- Chasing Consistency in Text-to-3D Generation from a Single Image
- Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
- Backdooring Textual Inversion for Concept Censorship
- DIFFNAT: Improving Diffusion Image Quality Using Natural Image Statistics
- HiFi Tuner: High-Fidelity Subject-Driven Fine-Tuning for Diffusion Models
- ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet
- DreaMoving: A Human Video Generation Framework based on Diffusion Models
- MotionCrafter: One-Shot Motion Customization of Diffusion Models
- Stellar: Systematic Evaluation of Human-Centric Personalized Text-to-Image Methods
- Towards Accurate Guided Diffusion Sampling through Symplectic Adjoint Method
- Object-Driven One-Shot Fine-tuning of Text-to-Image Diffusion with Prototypical Embedding
- StyleForge: Enhancing Text-to-Image Synthesis for Any Artistic Styles with Dual Binding
- SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
- Controllable Textual Inversion for Personalized Text-to-Image Generation
- Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation Knowledge
- SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
- VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models