activity
20202026
most citedReducing Language Biases in Visual Question Answering with Visually-Grounded Question Encoder

2 citations · 2 across the 9 of their papers we have counts for

collaborators

10 papers

cs.SD2026

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Ilpo Viertola, Giulio Cengarle, Gouthaman KV +2

We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the…

cs.SD2026

VIBE: Video Instruction-aligned Background music gEneration

Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj +5

Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representa…

cs.LG2026

A robust PPG foundation model using multimodal physiological supervision

Eloy Geenjaar, Vince Calhoun, Scott Daly +4

Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings. Recent PPG foundation models either us…

cs.SD2026

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5

Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…

cs.CV2025

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

Xiyang Wu, Zongxia Li, Jihui Jin +7

Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a nove…

cs.CV2025

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

Aakriti Agrawal, Gouthaman KV, Rohith Aralikatti +6

Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual information during multimodal reasoning.…