4 papers
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction
Ambuj Mehrish, Sebastiano Vascon
Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it produces realistic but stimulus-inaccura…
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
Ambuj Mehrish, Sebastiano Vascon
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority…
Constrained Dominant Sets for Multimodal Document Question Answering
Ambuj Mehrish, Sebastiano Vascon
Long multimodal document question answering is limited by which evidence reaches the reader, rather than by the quantity retrieved. In lengthy documents, findings often recur acros…
FLOWREADER: Min-Cost Flow Optimization for Multi-Modal Long Document Q&A
Ambuj Mehrish, Sebastiano Vascon
Long, multimodal documents force retrieval-augmented systems to assemble answers from evidence fragmented across text, tables, and slides broken across cells in a long table, sprea…