4 papers
Foundation Models are Implicit Deepfake Detectors
Stefan Smeu, Dragos-Alexandru Boldisor, Elisabeta Oneata +1
Pretrained self-supervised representations have emerged as a core component of current deepfake detection methods, yet it remains unclear which of their properties make real and fa…
Pathways of Visual Information Flow in Vision-Language Models
Israfel Salazar, Stella Frank, Dan Oneata +2
We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two dist…
Connecting Speech to Words through Images
Gabriel Pirlogeanu, Dan Oneata, Horia Cucu +1
How can we learn the mapping between written words and their spoken counterparts in the absence of explicit textual supervision? We present a visually grounded method for building…
Anchoring the Unknown: Open-Set Model Attribution via Proxy-Anchor Learning
Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu +3
The proliferation of text-to-speech (TTS) systems capable of generating realistic synthetic speech poses growing challenges for audio forensics. While binary deepfake detection has…