2 papers
cs.CL2025
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally +2
Offering a promising solution to the scalability challenges associated with human evaluation, the LLM-as-a-judge paradigm is rapidly gaining traction as an approach to evaluating l…
cs.LG2024
ICU-Sepsis: A Benchmark MDP Built from Real Medical Data
Kartik Choudhary, Dhawal Gupta, Philip S. Thomas
We present ICU-Sepsis, an environment that can be used in benchmarks for evaluating reinforcement learning (RL) algorithms. Sepsis management is a complex task that has been an imp…