11 papers
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
Tong Zhang, Motasem Alfarra, Carlos Hinojosa +2
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, furthe…
ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections
Kebin Contreras, Carlos Hinojosa, Jorge Bacca +1
Computer-use agents are increasingly capable of operating on real operating systems, but this capability has also increased the risks posed by prompt injection, indirect instructio…
CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation
Pablo Messina, Andrés Villa, Juan León Alcázar +5
Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign…
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
Abdelrahman Eldesokey, Merey Ramazanova, Ahmad Sait +4
Text-to-image (T2I) generation has advanced rapidly, making reliable evaluation critical as performance differences between models narrow. Existing evaluation practices typically a…
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
Ismael Elsharkawi, Ahmed Sait, Silvio Giancola +3
Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos due to large viewpoint variati…
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
Carlos Hinojosa, Clemens Grange, Bernard Ghanem
Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visua…