3 papers
cs.CV2025
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
Jinsheng Huang, Liang Chen, Taian Guo +13
Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, often assessed through multiple-choice questions (MCQs) that include an image,…
cs.CL2024
A Hybrid RAG System with Comprehensive Enhancement on Complex Reasoning
Ye Yuan, Chengwu Liu, Jingyang Yuan +3
Retrieval-augmented generation (RAG) is a framework enabling large language models (LLMs) to enhance their accuracy and reduce hallucinations by integrating external knowledge base…
cs.CL2024
Measuring Social Norms of Large Language Models
Ye Yuan, Kexin Tang, Jianhao Shen +2
We present a new challenge to examine whether large language models understand social norms. In contrast to existing datasets, our dataset requires a fundamental understanding of s…