2 papers
cs.CL2025
Training Report of TeleChat3-MoE
Xinzhang Liu, Chao Wang, Zhihao Yang +51
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one t…
cs.LG2025
Preventing Model Collapse via Contraction-Conditioned Neural Filters
Zongjian Han, Yiran Liang, Ruiwen Wang +4
This paper presents a neural network filter method based on contraction operators to address model collapse in recursive training of generative models. Unlike \cite{xu2024probabili…