4 papers
HADIS: Hybrid Adaptive Diffusion Model Serving for Efficient Text-to-Image Generation
Qizheng Yang, Tung-I Chen, Siyu Zhao +2
Text-to-image diffusion models have achieved remarkable visual quality but incur high computational costs, making latency-aware, scalable deployment challenging. To address this, w…
Unbiased Visual Reasoning with Controlled Visual Inputs
Zhaonan Li, Shijie Lu, Fei Wang +11
End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone whe…
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
Sohaib Ahmad, Qizheng Yang, Haoliang Wang +2
Text-to-image generation using diffusion models has gained increasing popularity due to their ability to produce high-quality, realistic images based on text prompts. However, effi…
AdapMTL: Adaptive Pruning Framework for Multitask Learning Model
Mingcan Xiang, Steven Jiaxun Tang, Qizheng Yang +2
In the domain of multimedia and multimodal processing, the efficient handling of diverse data streams such as images, video, and sensor data is paramount. Model compression and mul…