3 papers
cs.LG2026
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs +3
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these enviro…
cs.CL2025
Quantifying Phonosemantic Iconicity Distributionally in 6 Languages
George Flint, Kaustubh Kislay
Language is, as commonly theorized, largely arbitrary. Yet, systematic relationships between phonetics and semantics have been observed in many specific cases. To what degree could…
cs.LG2024
Evaluating K-Fold Cross Validation for Transformer Based Symbolic Regression Models
Kaustubh Kislay, Shlok Singh, Soham Joshi +4
Symbolic Regression remains an NP-Hard problem, with extensive research focusing on AI models for this task. Transformer models have shown promise in Symbolic Regression, but perfo…