3 papers
cs.CL2025
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
Taneesh Gupta, Shivam Shandilya, Xuchao Zhang +5
Reward modeling in large language models is susceptible to reward hacking, causing models to latch onto superficial features such as the tendency to generate lists or unnecessarily…
cs.CL2024
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
Shivam Shandilya, Menglin Xia, Supriyo Ghosh +4
The increasing prevalence of large language models (LLMs) such as GPT-4 in various applications has led to a surge in the size of prompts required for optimal performance, leading…
cs.LG2024
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
Aditya Soni, Mayukh Das, Anjaly Parayil +8
The difficulty of exploring and training online on real production systems limits the scope of real-time online data/feedback-driven decision making. The most feasible approach is…