2 papers
cs.CV2024
SegLLM: Multi-round Reasoning Segmentation
XuDong Wang, Shaolun Zhang, Shufan Li +5
We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual…
cs.CV2024
Aligning Diffusion Models by Optimizing Human Utility
Shufan Li, Konstantinos Kallidromitis, Akash Gokul +2
We present Diffusion-KTO, a novel approach for aligning text-to-image diffusion models by formulating the alignment objective as the maximization of expected human utility. Since t…