1 paper · 1 filter
De Jiang, Zhengyang Zhang, Kehong Yuan +1
Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy sourc…