Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Dual Attention Residuals
Xingda Yu, Yining Li, Xinzhang Liu +5
Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain…
cs.CL2025
Training Report of TeleChat3-MoE
Xinzhang Liu, Chao Wang, Zhihao Yang +51
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one t…
cs.CL2024★ 2 cited
TeleChat Technical Report
Zhongjiang He, Zihan Wang, Xinzhang Liu +33
In this technical report, we present TeleChat, a collection of large language models (LLMs) with parameters of 3 billion, 7 billion and 12 billion. It includes pretrained language…