1 citations · 1 across the 5 of their papers we have counts for
5 papers
Dual Attention Residuals
Xingda Yu, Yining Li, Xinzhang Liu +5
Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain…
Training Report of TeleChat3-MoE
Xinzhang Liu, Chao Wang, Zhihao Yang +51
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one t…
Technical Report of TeleChat2, TeleChat2.5 and T1
Zihan Wang, Xinzhang Liu, Yitong Yao +35
We introduce the latest series of TeleChat models: \textbf{TeleChat2}, \textbf{TeleChat2.5}, and \textbf{T1}, offering a significant upgrade over their predecessor, TeleChat. Despi…
52B to 1T: Lessons Learned via Tele-FLM Series
Xiang Li, Yiqun Yao, Xin Jiang +17
Large Language Models (LLMs) represent a significant stride toward Artificial General Intelligence. As scaling laws underscore the potential of increasing model sizes, the academic…
Tele-FLM Technical Report
Xiang Li, Yiqun Yao, Xin Jiang +17
Large language models (LLMs) have showcased profound capabilities in language understanding and generation, facilitating a wide array of applications. However, there is a notable p…