collaborators

7 papers

cs.AI2026

FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models

Bing Tian, Haikun Liu, Xiaocheng Zhong +5

Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks but making inter-block decoding strictly…

cs.AI2026

Akashic: A Low-Overhead LLM Inference Service with MemAttention

Yang Liu, Zhaokai Luo, Huayi Jin +7

Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every r…

cs.AI2026

RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

Yang Liu, ZhaoKai Luo, Zhaokai Luo +11

As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure. It limits GPU memory capacity, serv…

cs.IR2026

The Clustering Strikes Back: Building Cost-Effective and High-Performance ANNS at Scale with Helmsman

Yuchen Huang, Baiteng Ma, Yiping Sun +7

RedNote (a.k.a., Xiaohongshu, a global-scale social network platform) widely adopts approximate nearest neighbor search (ANNS) to power its search, recommendation, and advertising…

cs.IR2026

CCD-Level and Load-Aware Thread Orchestration for In-Memory Vector ANNS on Multi-Core CPUs

Yuchen Huang, Baiteng Ma, Yiping Sun +6

Vector approximate nearest neighbor search (ANNS) underpins search engines, recommendation systems, and advertising services. Recent advances in ANNS indexes make CPU a cost-effect…

cs.IR2026

LASER: An Efficient Target-Aware Segmented Attention Framework for End-to-End Long Sequence Modeling

Tianhe Lin, Ziwei Xiong, Baoyuan Ou +8

Modeling ultra-long user behavior sequences is pivotal for capturing evolving and lifelong interests in modern recommendation systems. However, deploying such models in real-time i…