14 papers
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction
Ambuj Mehrish, Sebastiano Vascon
Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it produces realistic but stimulus-inaccura…
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
Ambuj Mehrish, Sebastiano Vascon
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority…
Constrained Dominant Sets for Multimodal Document Question Answering
Ambuj Mehrish, Sebastiano Vascon
Long multimodal document question answering is limited by which evidence reaches the reader, rather than by the quantity retrieved. In lengthy documents, findings often recur acros…
FLOWREADER: Min-Cost Flow Optimization for Multi-Modal Long Document Q&A
Ambuj Mehrish, Sebastiano Vascon
Long, multimodal documents force retrieval-augmented systems to assemble answers from evidence fragmented across text, tables, and slides broken across cells in a long table, sprea…
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
Jan Melechovsky, Ambuj Mehrish, Abhinaba Roy +1
Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when create…
DialogXpert: Driving Intelligent and Emotion-Aware Conversations through Online Value-Based Reinforcement Learning with LLM Priors
Tazeek Bin Abdur Rakib, Ambuj Mehrish, Lay-Ki Soon +2
Large-language-model (LLM) agents excel at reactive dialogue but struggle with proactive, goal-driven interactions due to myopic decoding and costly planning. We introduce DialogXp…