8 papers
Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization
Ming Sun, Kun Yuan
Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinato…
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
Mingfei Sun
Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the F…
Learning what to say and how precisely: Efficient Communication via Differentiable Discrete Communication Learning
Aditya Kapoor, Yash Bhisikar, Benjamin Freed +2
Effective communication in multi-agent reinforcement learning (MARL) is critical for success but constrained by bandwidth, yet past approaches have been limited to complex gating m…
Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning
Aditya Kapoor, Kale-ab Tessera, Mayank Baranwal +4
Credit assignmen, disentangling each agent's contribution to a shared reward, is a critical challenge in cooperative multi-agent reinforcement learning (MARL). To be effective, cre…
Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem
Guorui Quan, Mingfei Sun, Manuel López-Ibáñez
The art of heuristic design has traditionally been a human pursuit. While Large Language Models (LLMs) can generate code for search heuristics, their application has largely been c…
Generating Query-Relevant Document Summaries via Reinforcement Learning
Nitin Yadav, Changsung Kang, Hongwei Shang +1
E-commerce search engines often rely solely on product titles as input for ranking models with latency constraints. However, this approach can result in suboptimal relevance predic…