2 citations · 3 across the 6 of their papers we have counts for
4 papers · 1 filter
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
Yang Zhang, Danyang Li, Yuxuan Li +4
Multimodal Large Language Models (MLLMs) have achieved remarkable performance by integrating powerful language backbones with large-scale visual encoders. Among these, latent Chain…
Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized
Er Jin, Yang Zhang, Yongli Mou +4
Recent advances in generative models have demonstrated an exceptional ability to produce highly realistic images. However, previous studies show that generated images often resembl…
Learnable Sparsity for Vision Generative Models
Yang Zhang, Er Jin, Wenzhong Liang +5
Diffusion models have achieved impressive advancements in various vision tasks. However, these gains often rely on increasing model size, which escalates computational complexity a…
Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models
Yang Zhang, Teoh Tze Tzun, Lim Wei Hern +2
Recent advancements in diffusion models have notably improved the perceptual quality of generated images in text-to-image synthesis tasks. However, diffusion models often struggle…