34 citations · 57 across the 15 of their papers we have counts for
7 papers · 1 filter
Sigma-MoE-Tiny Technical Report
Qingguo Hu, Zhenghao Lin, Ziyue Yang +12
Mixture-of-Experts (MoE) has emerged as a promising paradigm for foundation models due to its efficient and powerful scalability. In this work, we present Sigma-MoE-Tiny, an MoE la…
SIGMA: An AI-Empowered Training Stack on Early-Life Hardware
Lei Qu, Lianhai Ren, Peng Cheng +12
An increasing variety of AI accelerators is being considered for large-scale training. However, enabling large-scale training on early-life AI accelerators faces three core challen…
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
Haoran Deng, Yingyu Lin, Zhenghao Lin +4
Long-context language models unlock advanced capabilities in reasoning, code generation, and document summarization by leveraging dependencies across extended spans of text. Howeve…
Learning from the Best, Differently: A Diversity-Driven Rethinking on Data Selection
Hongyi He, Xiao Liu, Zhenghao Lin +6
High-quality pre-training data is crutial for large language models, where quality captures factual reliability and semantic value, and diversity ensures broad coverage and distrib…
LayerNorm Induces Recency Bias in Transformer Decoders
Junu Kim, Xiao Liu, Zhenghao Lin +3
Causal self-attention provides positional information to Transformer decoders. Prior work has shown that stacks of causal self-attention layers alone induce a positional bias in at…
Generalized Category Discovery in Event-Centric Contexts: Latent Pattern Mining with LLMs
Yi Luo, Qiwen Wang, Junqi Yang +5
Generalized Category Discovery (GCD) aims to classify both known and novel categories using partially labeled data that contains only known classes. Despite achieving strong perfor…