works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.AR2026

Realizable N:M Sparse Transformer Inference via Search-Kernel Co-Design

Yiming Liu, Wenqi Lou, Zhiguang Wang +4

The paper presents a co-designed hardware and software framework that enables fast inference of Vision Transformers by applying N:M structured sparsity with a specialized CUDA kern…

cs.IR2026

OneRanker: Unified Generation and Ranking with One Model in Industrial Advertising Recommendation

Dekai Sun, Yiming Liu, Jiafan Zhou +6

The end-to-end generative paradigm is revolutionizing advertising recommendation systems, driving a shift from traditional cascaded architectures towards unified modeling. However,…

cs.DC2026

xLLM Technical Report

Tongxuan Liu, Tao Peng, Peijun Yang +50

We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimi…

cs.LG2026

Window-Diffusion: Accelerating Diffusion Language Model Inference with Windowed Token Pruning and Caching

Fengrui Zuo, Zhiwei Ke, Yiming Liu +3

Diffusion language models (DLMs) generate text through iterative denoising, but inference requires full-sequence attention at every iteration, resulting in substantial redundant co…

cs.LG2026

AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving

Tianhao Xu, Yiming Liu, Xianglong Lu +18

Optimizing Large Language Model (LLM) inference in production systems is increasingly difficult due to dynamic workloads, stringent latency/throughput targets, and a rapidly expand…

cs.CR2025

Can Watermarked LLMs be Identified by Users via Crafted Prompts?

Aiwei Liu, Sheng Guan, Yiming Liu +6

Text watermarking for Large Language Models (LLMs) has made significant progress in detecting LLM outputs and preventing misuse. Current watermarking techniques offer high detectab…