7 papers
Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
Aakriti Agrawal, Gouthaman KV, Rohith Aralikatti +6
Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual information during multimodal reasoning.…
A robust PPG foundation model using multimodal physiological supervision
Eloy Geenjaar, Vince Calhoun, Scott Daly +4
Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings. Recent PPG foundation models either us…
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
Vaibhavi Lokegaonkar, Aryan Vijay Bhosale, Vishnu Raj +5
Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typically relying on visual conditi…
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
Xiyang Wu, Zongxia Li, Jihui Jin +7
Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a nove…
Enhanced Generative Machine Listener
Vishnu Raj, Gouthaman KV, Shiv Gehlot +2
We present GMLv2, a reference-based model designed for the prediction of subjective audio quality as measured by MUSHRA scores. GMLv2 introduces a Beta distribution-based loss to m…
Moment Sampling in Video LLMs for Long-Form Video QA
Mustafa Chasmai, Gauri Jagatap, Gouthaman KV +3
Recent advancements in video large language models (Video LLMs) have significantly advanced the field of video question answering (VideoQA). While existing methods perform well on…