2 citations · 2 across the 10 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
EchoRL: Reinforcement Learning via Rollout Echoing
Jinhe Bi, Aniri, Minglai Yang +9
Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models. However, as training proceed…
cs.LG2026
Memory-R2: Fair Credit Assignment for Long-Horizon Memory-Augmented LLM Agents
Sikuan Yan, Ahmed Bahloul, Ercong Nie +4
Memory-augmented LLM agents enable interactions that extend beyond finite context windows by storing, updating, and reusing information across sessions. However, training such agen…
cs.LG2026
Routing-Free Mixture-of-Experts
Yilun Liu, Jinru Han, Sikuan Yan +2
Standard Mixture-of-Experts (MoE) models rely on centralized routing mechanisms that introduce rigid inductive biases. We propose Routing-Free MoE which eliminates any hard-coded c…