5 papers
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning in Agents
Neharika Jali, Anupam Nayak, Gauri Joshi
As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces even for simple queries. Prior appro…
ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
Shivam Patel, Neharika Jali, Ankur Mallick +1
Large language model (LLM) query routers are critical to modern AI platforms as they seek to improve efficiency by assigning inference queries to accurate, yet low-cost models. Par…
Natural Policy Gradient for Average Reward Non-Stationary RL
Neharika Jali, Eshika Pathak, Pranay Sharma +2
We consider the problem of non-stationary reinforcement learning (RL) in the infinite-horizon average-reward setting. We model it by a Markov Decision Process with time-varying rew…
Erasure Coded Neural Network Inference via Fisher Averaging
Divyansh Jhunjhunwala, Neharika Jali, Gauri Joshi +1
Erasure-coded computing has been successfully used in cloud systems to reduce tail latency caused by factors such as straggling servers and heterogeneous traffic variations. A majo…
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
Neharika Jali, Guannan Qu, Weina Wang +1
We consider the problem of efficiently routing jobs that arrive into a central queue to a system of heterogeneous servers. Unlike homogeneous systems, a threshold policy, that rout…