8 citations · 14 across the 17 of their papers we have counts for
3 papers · 1 filter
Can Sound Replace Vision in LLaVA With Token Substitution?
Ali Vosoughi, Jing Bi, Pinxin Liu +2
What happens when we push audio-visual alignment to its absolute limits? To systematically investigate this question, we needed datasets with granular alignment quality annotations…
Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model
Ali Vosoughi, Dimitra Emmanouilidou, Hannes Gamper
Integrating audio and visual data for training multimodal foundational models remains a challenge. The Audio-Video Vector Alignment (AVVA) framework addresses this by considering A…
Learning Audio Concepts from Counterfactual Natural Language
Ali Vosoughi, Luca Bondi, Ho-Hsiang Wu +1
Conventional audio classification relied on predefined classes, lacking the ability to learn from free-form text. Recent methods unlock learning joint audio-text embeddings from ra…