3 papers
cs.SD2026
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
Gokul Karthik Kumar, Ludovick Lepauloux, Hakim Hacid
Whisper has become the de-facto encoder for extracting general-purpose audio features in large audio-language models, where a 30-second clip is typically represented by 1500 frame…
cs.SD2026
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
Gokul Karthik Kumar, Rishabh Saraf, Ludovick Lepauloux +3
Large language models (LLMs) have transformed NLP, yet their integration with audio remains underexplored despite audio's centrality to human communication. We introduce Falcon3-Au…
cs.CL2025
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
Gokul Karthik Kumar, Iheb Chaabane, Kebin Wu
Vision-language models (VLMs) excel in various visual benchmarks but are often constrained by the lack of high-quality visual fine-tuning data. To address this challenge, we introd…