2 papers
cs.LG2026
When Less is Enough: Efficient Inference via Collaborative Reasoning
Yilei Chen, Sharut Gupta, Yannis Paschalidis +2
In this work, we introduce DUET (Dual-model Efficient Two-stage inference), a collaborative inference framework in which a capable model and a lightweight model work together to so…
cs.CL2026
Post-training Large Language Models for Diverse High-Quality Responses
Yilei Chen, Souradip Chakraborty, Lorenz Wolf +2
Reinforcement learning (RL) has emerged as a popular method for post-training large language models (LLMs). While improving the model's performance on downstream tasks, it often re…