activity
20152026
most citedA robust and efficient video representation for action recognition

17 citations · 37 across the 23 of their papers we have counts for

collaborators
Showing 2025Show all

6 papers · 1 filter

cs.CV2025

Investigating self-supervised representations for audio-visual deepfake detection

Dragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata +1

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses th…

cs.CV2025

Not All Splits Are Equal: Rethinking Attribute Generalization Across Unrelated Categories

Liviu Nicolae Fircă, Antonio Bărbălau, Dan Oneata +1

Can models generalize attribute knowledge across semantically and perceptually dissimilar categories? While prior work has addressed attribute prediction within narrow taxonomic or…

cs.CL2025

The mutual exclusivity bias of bilingual visually grounded speech models

Dan Oneata, Leanne Nortje, Yevgen Matusevych +1

Mutual exclusivity (ME) is a strategy where a novel word is associated with a novel object rather than a familiar one, facilitating language learning in children. Recent work has f…

cs.CL2025

Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era

Dan Oneata, Desmond Elliott, Stella Frank

Human learning and conceptual representation is grounded in sensorimotor experience, in contrast to state-of-the-art foundation models. In this paper, we investigate how well such…

eess.AS20256 cited

Unmasking real-world audio deepfakes: A data-centric approach

David Combei, Adriana Stan, Dan Oneata +2

The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific pur…

eess.AS20254 cited

TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes

Adriana Stan, David Combei, Dan Oneata +1

Deepfake detection has gained significant attention across audio, text, and image modalities, with high accuracy in distinguishing real from fake. However, identifying the exact so…