2 papers
cs.CV2024
Meerkat: Audio-Visual Large Language Model for Grounding in Space and Time
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta +4
Leveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. Howeve…
cs.CV2023
AdVerb: Visually Guided Audio Dereverberation
Sanjoy Chowdhury, Sreyan Ghosh, Subhrajyoti Dasgupta +3
We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverber…