3 papers
cs.LG2026
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these enviro…
cs.CL2025
When Two LLMs Debate, Both Think They'll Win
Pradyumna Shyama Prasad, Minh Nhat Nguyen
Can LLMs accurately adjust their confidence when facing opposition? Building on previous studies measuring calibration on static fact-based question-answering tasks, we evaluate La…
cs.IR2025
Know Or Not: a library for evaluating out-of-knowledge base robustness
Jessica Foo, Pradyumna Shyama Prasad, Shaun Khoo
While the capabilities of large language models (LLMs) have progressed significantly, their use in high-stakes applications have been limited due to risks of hallucination. One key…