3 papers
cs.AI2026
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
Wenjing Zhang, Jiangze Yan, Jieyun Huang +7
Distilling reasoning capabilities from Large Reasoning Models (LRMs) into smaller models is typically constrained by the limitation of rejection sampling. Standard methods treat th…
cs.CL2026
LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Experts
Yuan Zhuang, Yi Shen, Yuexin Bian +4
Recent studies have shown that combining parameter-efficient fine-tuning (PEFT) with mixture-of-experts (MoE) is an effective strategy for adapting large language models (LLMs) to…
cs.MA2025
YOLO-MARL: You Only LLM Once for Multi-Agent Reinforcement Learning
Yuan Zhuang, Yi Shen, Zhili Zhang +2
Advancements in deep multi-agent reinforcement learning (MARL) have positioned it as a promising approach for decision-making in cooperative games. However, it still remains challe…