5 papers
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
Guanhua Chen, Chuyue Huang, Yutong Yao +4
Multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities (entire images or scenes), creating a mismatch with fine-grained user queries and…
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
Guanhua Chen, Yutong Yao, Shenghe Sun +5
Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure question answering (VP-QA)…
SGIC: A Self-Guided Iterative Calibration Framework for RAG
Guanhua Chen, Yutong Yao, Lidia S. Chao +2
Recent research in retrieval-augmented generation (RAG) has concentrated on retrieving useful information from candidate documents. However, numerous methodologies frequently negle…
Not All LoRA Parameters Are Essential: Insights on Inference Necessity
Guanhua Chen, Yutong Yao, Ci-Jun Gao +3
Current research on LoRA primarily focuses on minimizing the number of fine-tuned parameters or optimizing its architecture. However, the necessity of all fine-tuned LoRA layers du…
PMMT: Preference Alignment in Multilingual Machine Translation via LLM Distillation
Shuqiao Sun, Yutong Yao, Peiwen Wu +2
Translation is important for cross-language communication, and many efforts have been made to improve its accuracy. However, less investment is conducted in aligning translations w…