3 papers
cs.MA2026
Adaptive Value Decomposition: Coordinating a Varying Number of Agents in Urban Systems
Yexin Li, Jinjin Guo, Haoyu Zhang +3
Multi-agent reinforcement learning (MARL) provides a promising paradigm for coordinating multi-agent systems (MAS). However, most existing methods rely on restrictive assumptions,…
cs.LG2026
Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning
Jinjin Guo, Yexin Li, Zhichao Huang +5
Large-scale multimodal contrastive learning has recently achieved impressive success in learning rich and transferable representations, yet it remains fundamentally limited by the…
cs.LG2026
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
Zheng Zhang, Ao Lu, Yuanhao Zeng +5
Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant breakthroughs in complex LLM reasoning within verifiable domains, such as mathematics and programmin…