2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CL2025
Mitigating Self-Preference by Authorship Obfuscation
Taslim Mahbub, Shi Feng
Language models (LMs) judges are widely used to evaluate the quality of LM outputs. Despite many advantages, LM judges display concerning biases that can impair their integrity in…
cs.LG2025★ 2 cited
Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
Khizar Anjum, Muhammad Arbab Arshad, Kadhim Hayawi +10
Large language models (LLMs) are increasingly being deployed across disciplines due to their advanced reasoning and problem solving capabilities. To measure their effectiveness, va…