3 papers
cs.DB2026
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
Jianxin Yan, Zeheng Qian, Wangze Ni +6
Cache fusion accelerates generation process of LLMs equipped with RAG through KV caching and selective token recomputation, thereby reducing computational costs and improving effic…
cs.IR2026
SRBench: A Comprehensive Benchmark for Sequential Recommendation with Large Language Models
Jianhong Li, Zeheng Qian, Wangze Ni +4
LLM development has aroused great interest in Sequential Recommendation (SR) applications. However, comprehensive evaluation of SR models remains lacking due to the limitations of…
cs.CR2026
SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails
Zhiyi Mou, Jingyuan Yang, Zeheng Qian +6
While Large Language Models (LLMs) have powerful capabilities, they remain vulnerable to jailbreak attacks, which is a critical barrier to their safe web real-time application. Cur…