Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
Martino M. L. Pulici, Cuong Xuan Chu, Evgeny Kharlamov +3
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models (…
cs.LG2025
Distilling Tool Knowledge into Language Models via Back-Translated Traces
Xingyue Huang, Xianglong Hu, Zifeng Ding +9
Large language models (LLMs) often struggle with mathematical problems that require exact computation or multi-step algebraic reasoning. Tool-integrated reasoning (TIR) offers a pr…