activity
20242026
collaborators

8 papers

cs.LG2026

Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization

Ming Sun, Kun Yuan

Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinato…

cs.LG2026

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation

Mingfei Sun

Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the F…

cs.MA2025

Learning what to say and how precisely: Efficient Communication via Differentiable Discrete Communication Learning

Aditya Kapoor, Yash Bhisikar, Benjamin Freed +2

Effective communication in multi-agent reinforcement learning (MARL) is critical for success but constrained by bandwidth, yet past approaches have been limited to complex gating m…

cs.MA2025

Redistributing Rewards Across Time and Agents for Multi-Agent Reinforcement Learning

Aditya Kapoor, Kale-ab Tessera, Mayank Baranwal +4

Credit assignmen, disentangling each agent's contribution to a shared reward, is a critical challenge in cooperative multi-agent reinforcement learning (MARL). To be effective, cre…

cs.AI2025

Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem

Guorui Quan, Mingfei Sun, Manuel López-Ibáñez

The art of heuristic design has traditionally been a human pursuit. While Large Language Models (LLMs) can generate code for search heuristics, their application has largely been c…

cs.IR2025

Generating Query-Relevant Document Summaries via Reinforcement Learning

Nitin Yadav, Changsung Kang, Hongwei Shang +1

E-commerce search engines often rely solely on product titles as input for ranking models with latency constraints. However, this approach can result in suboptimal relevance predic…