works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CR2026

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

Yi Li, Cheng Li, Chen Li +2

The paper proposes a privacy‑focused edge‑cloud collaborative framework for large language model inference that authenticates and encrypts KV cache data, allowing lightweight edge…

cs.LG2025

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts

Cheng Li, Jiexiong Liu, Yixuan Chen +1

Transformer models based on the Mixture of Experts (MoE) architecture have made significant progress in long-sequence modeling, but existing models still have shortcomings in compu…

cs.AI2025

KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models

Cheng Li, Jiexiong Liu, Yixuan Chen +2

This paper introduces KunLunBaizeRAG, a reinforcement learning-driven reasoning framework designed to enhance the reasoning capabilities of large language models (LLMs) in complex…

cs.AI2025

Video-VoT-R1: An efficient video inference model integrating image packing and AoE architecture

Cheng Li, Jiexiong Liu, Yixuan Chen +1

In the field of video-language pretraining, existing models face numerous challenges in terms of inference efficiency and multimodal data processing. This paper proposes a KunLunBa…

cs.CL2025

KunlunBaize: LLM with Multi-Scale Convolution and Multi-Token Prediction Under TransformerX Framework

Cheng Li, Jiexiong Liu, Yixuan Chen +2

Large language models have demonstrated remarkable performance across various tasks, yet they face challenges such as low computational efficiency, gradient vanishing, and difficul…