31 citations · 70 across the 25 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Explicit Reasoning Makes Better Judges: A Systematic Study on Accuracy, Efficiency, and Robustness
Pratik Jayarao, Himanshu Gupta, Neeraj Varshney +1
As Large Language Models (LLMs) are increasingly adopted as automated judges in benchmarking and reward modeling, ensuring their reliability, efficiency, and robustness has become…
cs.AI2023
Methods and Mechanisms for Interactive Novelty Handling in Adversarial Environments
Tung Thai, Ming Shen, Mayank Garg +10
Learning to detect, characterize and accommodate novelties is a challenge that agents operating in open-world domains need to address to be able to guarantee satisfactory task perf…