5 papers
A Unified Zeroth-Order Approach for Decentralized Minimax Optimization
Haoyuan Cai, Yike Zhao, Aleksandar Armacki +2
We propose ZOMA, a unified Zeroth-Order decentralized accelerated MinimAx framework for multi-agent nonconvex Polyak--Åojasiewicz minimax optimization. The proposed framework only…
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
Yike Zhao, Onno Eberhard, Malek Khammassi +2
The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning. We provide a theoretical justi…
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
Xinyuan Li, Murong Xu, Wenbiao Tao +4
Large language models (LLMs) achieve high performance on mathematical reasoning, but these results can be inflated by training data leakage or superficial pattern matching rather t…
More Data or Better Data? A Critical Analysis of Data Selection and Synthesis for Mathematical Reasoning
Yike Zhao, Simin Guo, Ziqing Yang +3
The reasoning capabilities of Large Language Models (LLMs) play a critical role in many downstream tasks, yet depend strongly on the quality of training data. Despite various propo…
Diffusion Stochastic Learning Over Adaptive Competing Networks
Yike Zhao, Haoyuan Cai, Ali H. Sayed
This paper studies a stochastic dynamic game between two competing teams, each consisting of a network of collaborating agents. Unlike fully cooperative settings, where all agents…