1 citations · 1 across the 6 of their papers we have counts for
6 papers
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5
Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
Meng Chen, Kun Wang, Li Lu +2
Modern Large audio-language models (LALMs) power intelligent voice interactions by tightly integrating audio and text. This integration, however, expands the attack surface beyond…
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
S Sakshi, Vaibhavi Lokegaonkar, Neil Zhang +4
Spatial perception is central to auditory intelligence, enabling accurate understanding of real-world acoustic scenes and advancing human-level perception of the world around us. W…
Transformation of audio embeddings into interpretable, concept-based representations
Alice Zhang, Edison Thomaz, Lie Lu
Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to in…
Semi-Supervised Contrastive Learning for Controllable Video-to-Music Retrieval
Shanti Stewart, Gouthaman KV, Lie Lu +1
Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best musi…
CiTrus: Squeezing Extra Performance out of Low-data Bio-signal Transfer Learning
Eloy Geenjaar, Lie Lu
Transfer learning for bio-signals has recently become an important technique to improve prediction performance on downstream tasks with small bio-signal datasets. Recent works have…