2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CL2026
GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning
Lam So, Canhui Wu, Han Lin
Reasoning-oriented large language models often achieve strong problem-solving performance by generating long chains of thought, but this behavior substantially increases inference…
cs.AI2026★ 2 cited
RUMAD: Reinforcement-Unifying Multi-Agent Debate
Chao Wang, Han Lin, Huaze Tang +2
Multi-agent debate (MAD) systems leverage collective intelligence to enhance reasoning capabilities, yet existing approaches struggle to simultaneously optimize accuracy, consensus…