2 papers
cs.HC2026
Calibrating LLM Judges for Human and AI Conversations
Maike Züfle, Patrícia Schmidtová, Vilém Zouhar +8
Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of convers…
cs.CL2026
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
Kaushal Bhogale, Manas Dhir, Amritansh Walecha +11
Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference…