5 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Robust Reward Modeling via Causal Rubrics
Pragya Srivastava, Harman Singh, Rahul Madhavan +9
Reward models (RMs) are fundamental to aligning Large Language Models (LLMs) via human feedback, yet they often suffer from reward hacking. They tend to latch on to superficial or…
cs.CL2022★ 5 cited
CST5: Data Augmentation for Code-Switched Semantic Parsing
Anmol Agarwal, Jigar Gupta, Rahul Goel +3
Extending semantic parsers to code-switched input has been a challenging problem, primarily due to a lack of supervised training data. In this work, we introduce CST5, a new data a…