5 papers
Robust Sequential Experimental Design for A/B Testing
Qianglin Wen, Xiangkun Wu, Chengchun Shi +4
Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We st…
Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
Xiangxu Zhang, Lei Li, Yanyun Zhou +3
Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLMs) remain limited in capturing k…
Copy-Paste to Mitigate Large Language Model Hallucinations
Yongchao Long, Xian Wu, Yingying Zhang +3
While Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to generate contextually grounded responses, contextual faithfulness remains challenging as LLMs may…
Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
Jinlin Li, Yuran Wang, Yifei Yuan +5
Large Vision-Language Models (LVLMs) have recently achieved impressive results in multimodal tasks such as image captioning and visual question answering. However, they remain pron…
From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering
Lei Li, Xiao Zhou, Yingying Zhang +1
Medical question answering (QA) requires extensive access to domain-specific knowledge. A promising direction is to enhance large language models (LLMs) with external knowledge ret…