From the 1 of 7 linked papers with an AI index.
4 papers · 1 filter
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7
LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +4
The predominant approach to advancing text-to-image generation has been training-time scaling, where larger models are trained on more data using greater computational resources. W…
SegLLM: Multi-round Reasoning Segmentation
XuDong Wang, Shaolun Zhang, Shufan Li +5
We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual…
Aligning Diffusion Models by Optimizing Human Utility
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2
We present Diffusion-KTO, a novel approach for aligning text-to-image diffusion models by formulating the alignment objective as the maximization of expected human utility. Since t…