From the 1 of 5 linked papers with an AI index.
5 papers
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
Yi Li, Cheng Li, Chen Li +2
The paper proposes a privacy‑focused edge‑cloud collaborative framework for large language model inference that authenticates and encrypts KV cache data, allowing lightweight edge…
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts
Cheng Li, Jiexiong Liu, Yixuan Chen +1
Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in compu…
KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
Cheng Li, Jiexiong Liu, Yixuan Chen +2
This paper introduces KunLunBaizeRAG, a reinforcement learning-driven reasoning framework designed to enhance the reasoning capabilities of large language models (LLMs) in complex…
Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture
Cheng Li, Jiexiong Liu, Yixuan Chen +1
In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBa…
KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework
Cheng Li, Jiexiong Liu, Yixuan Chen +2
Large language models have demonstrated remarkable performance across various tasks, yet they face challenges such as low computational efficiency, gradient vanishing, and difficul…