12 papers
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
Tong Zhang, Motasem Alfarra, Carlos Hinojosa +2
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, furthe…
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models
Aznaur Aliev, Carlos Hinojosa, Abdelrahman Eldesokey +3
HyperSafe introduces a post‑hoc, model‑specific safe side network generated by a hypernetwork that classifies prompts using activation fingerprints, allowing fine‑tuned language mo…
Defending Against Harmful Supervision Hidden in Benign Samples
Bang An, Yibo Yang, Dandan Guo +3
Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign ta…
ProjGuard: Safety Monitoring for Computer-Use Agents via Low-Dimensional Projections
Kebin Contreras, Carlos Hinojosa, Jorge Bacca +1
Computer-use agents are increasingly capable of operating on real operating systems, but this capability has also increased the risks posed by prompt injection, indirect instructio…
CURE: Curriculum-guided Multi-task Training for Reliable Anatomy Grounded Report Generation
Pablo Messina, Andrés Villa, Juan León Alcázar +5
Medical vision-language models can automate the generation of radiology reports but struggle with accurate visual grounding and factual consistency. Existing models often misalign…
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
Carlos Hinojosa, Clemens Grange, Bernard Ghanem
Vision-language models (VLMs) are increasingly deployed in real-world and embodied settings where safety decisions depend on visual context. However, it remains unclear which visua…