Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
Gemma Zhang, Prachi Badarayani, Asmi Kumar +2
LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has…
cs.LG2026
Metag: A dataset to build agentic meta-reviewing capabilities
Anirudh Sundar, Min Chen, Divya Tadimeti +11
AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same time, the continuing growth…