3 papers
cs.SD2026
Scaling Audio-Text Retrieval with Multimodal Large Language Models
Jilan Xu, Carl Thomé, Danijela Horak +2
Audio-text retrieval is crucial for bridging acoustic signals and natural language. While contrastive dual-encoder architectures like CLAP have shown promise, they are fundamentall…
cs.CV2025
Towards Reliable Identification of Diffusion-based Image Manipulations
Alex Costanzino, Woody Bayliss, Juil Sock +5
Changing facial expressions, gestures, or background details may dramatically alter the meaning conveyed by an image. Notably, recent advances in diffusion models greatly improve t…
cs.AI2025
Intent Factored Generation: Unleashing the Diversity in Your Language Model
Eltayeb Ahmed, Uljad Berdica, Martha Elliott +2
Obtaining multiple meaningfully diverse, high quality samples from Large Language Models for a fixed prompt remains an open challenge. Current methods for increasing diversity ofte…