5 papers
Early-stopping for Transformer model training
Jing He, Hua Jiang, Cheng Li +2
This work, based on Random Matrix Theory (RMT), introduces a novel early-stopping strategy for Transformer training dynamics. Utilizing the Power Law (PL) fit to tansformer attenti…
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
Cheng Li, Jiexiong Liu, Yixuan Chen +1
Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in compu…
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
Cheng Li, Jiexiong Liu, Yixuan Chen +2
This paper introduces KunLunBaizeRAG, a reinforcement learning-driven reasoning framework designed to enhance the reasoning capabilities of large language models (LLMs) in complex…
Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
Cheng Li, Jiexiong Liu, Yixuan Chen +1
In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBa…
KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
Cheng Li, Jiexiong Liu, Yixuan Chen +2
Large language models have demonstrated remarkable performance across various tasks, yet they face challenges such as low computational efficiency, gradient vanishing, and difficul…