33 citations · 35 across the 9 of their papers we have counts for
14 papers · 1 filter
Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
Mingqi Gao, Anthony Sicilia, Weiyan Shi
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inferenc…
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
Junzhe Zhang, Huixuan Zhang, Xinyu Hu +4
Evaluation is important for multimodal generation tasks, while traditional multimodal evaluation metrics suffer from several limitations. With the rapid progress of MLLMs, there is…
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
Jiayi Chang, Mingqi Gao, Xinyu Hu +1
Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capab…
MMCR: Benchmarking Cross-Source Reasoning in Scientific Papers
Yang Tian, Zheng Lu, Mingqi Gao +2
Fully comprehending scientific papers by machines reflects a high level of Artificial General Intelligence, requiring the ability to reason across fragmented and heterogeneous sour…
Aspect-Guided Multi-Level Perturbation Analysis of Large Language Models in Automated Peer Review
Jiatao Li, Yanheng Li, Xinyu Hu +2
We propose an aspect-guided, multi-level perturbation framework to evaluate the robustness of Large Language Models (LLMs) in automated peer review. Our framework explores perturba…
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
Mingqi Gao, Yixin Liu, Xinyu Hu +3
Evaluating and ranking the capabilities of different LLMs is crucial for understanding their performance and alignment with human preferences. Due to the high cost and time-consumi…