3 papers
cs.CV2026
EDITOR: Effective and Interpretable Prompt Inversion for Text-to-Image Diffusion Models
Mingzhe Li, Kejing Xia, Gehao Zhang +5
Text-to-image generation models~(e.g., Stable Diffusion) have achieved significant advancements, enabling the creation of high-quality and realistic images based on textual descrip…
cs.SD2026
AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs
Townim Faisal Chowdhury, Ta Duc Huy, Siqi Pan +2
Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability i…
cs.CV2025
PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
Mingzhe Li, Renhao Zhang, Zhiyang Wen +4
Text-to-image (T2I) generative models such as Stable Diffusion and FLUX can synthesize realistic, high-quality images directly from textual prompts. The resulting image quality dep…