3 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 2 cited
Unlocking the Potential of Text-to-Image Diffusion with PAC-Bayesian Theory
Eric Hanchen Jiang, Yasi Zhang, Zhi Zhang +4
Text-to-image (T2I) diffusion models have revolutionized generative modeling by producing high-fidelity, diverse, and visually realistic images from textual prompts. Despite these…
cs.CV2024
SegLLM: Multi-round Reasoning Segmentation
XuDong Wang, Shaolun Zhang, Shufan Li +5
We present SegLLM, a novel multi-round interactive reasoning segmentation model that enhances LLM-based segmentation by exploiting conversational memory of both visual and textual…
cs.CV2021★ 3 cited
Interpreting Audiograms with Multi-stage Neural Networks
Shufan Li, Congxi Lu, Linkai Li +3
Audiograms are a particular type of line charts representing individuals' hearing level at various frequencies. They are used by audiologists to diagnose hearing loss, and further…