Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
dots.llm1 Technical Report
Bi Huo, Bin Tu, Cheng Qin +24
Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this…
cs.CL2024
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
Junfeng Tian, Da Zheng, Yang Cheng +3
Large language models (LLM) have prioritized expanding the context window from which models can incorporate more information. However, training models to handle long contexts prese…
cs.CL2024
Nyonic Technical Report
Junfeng Tian, Rui Wang, Cong Li +3
This report details the development and key achievements of our latest language model designed for custom large language models. The advancements introduced include a novel Online…