3 papers
cs.CL2026
SpecFuse: Ensembling Large Language Models via Next-Segment Prediction
Bo Lv, Nayu Liu, Chen Tang +3
Ensembles of generative large language models (LLMs) are a promising way to compensate for individual model limitations, integrating the strengths of different LLMs. Existing LLM e…
cs.CL2026
GRAPHMOE: Amplifying Cognitive Depth of Mixture-of-Experts Network via Introducing Self-Rethinking Mechanism
Bo Lv, Chen Tang, Zifan Zheng +8
Traditional Mixture-of-Experts (MoE) networks benefit from utilizing multiple smaller expert models as opposed to a single large network. However, these experts typically operate i…
cs.CL2025
HyCoRA: Hyper-Contrastive Role-Adaptive Learning for Role-Playing
Shihao Yang, Zhicong Lu, Yong Yang +3
Multi-character role-playing aims to equip models with the capability to simulate diverse roles. Existing methods either use one shared parameterized module across all roles or ass…