6 papers · 1 filter
Dual Attention Residuals
Xingda Yu, Yining Li, Xinzhang Liu +5
Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain…
Training Report of TeleChat3-MoE
Xinzhang Liu, Chao Wang, Zhihao Yang +51
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one t…
Technical Report of TeleChat2, TeleChat2.5 and T1
Zihan Wang, Xinzhang Liu, Yitong Yao +35
We introduce the latest series of TeleChat models: \textbf{TeleChat2}, \textbf{TeleChat2.5}, and \textbf{T1}, offering a significant upgrade over their predecessor, TeleChat. Despi…
52B to 1T: Lessons Learned via Tele-FLM Series
Xiang Li, Yiqun Yao, Xin Jiang +17
Large Language Models (LLMs) represent a significant stride toward Artificial General Intelligence. As scaling laws underscore the potential of increasing model sizes, the academic…
Tele-FLM Technical Report
Xiang Li, Yiqun Yao, Xin Jiang +17
Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable p…
TeleChat Technical Report
Zhongjiang He, Zihan Wang, Xinzhang Liu +33
In this technical report, we present TeleChat, a collection of large language models (LLMs) with parameters of 3 billion, 7 billion and 12 billion. It includes pretrained language…