kl regularization 1language model reasoning 1policy optimization 1reinforcement learning 1self-distillation 1token-level mixing 1
From the 1 of 10 linked papers with an AI index.
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Image Generation with a Sphere Encoder
Kaiyu Yue, Menglin Jia, Ji Hou +1
We introduce the Sphere Encoder, an efficient generative framework capable of producing images in a single forward pass and competing with many-step diffusion models using fewer th…
cs.CV2025
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
Kevin David Hayes, Micah Goldblum, Vikash Sehwag +3
Text-to-image (T2I) models are capable of generating visually impressive images, yet they often fail to accurately capture specific attributes in user prompts, such as the correct…
cs.CV2025
Analysis of Attention in Video Diffusion Transformers
Yuxin Wen, Jim Wu, Ajay Jain +2
We conduct an in-depth analysis of attention in video diffusion transformers (VDiTs) and report a number of novel findings. We identify three key properties of attention in VDiTs:…