3 papers
cs.DB2026
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
Jianxin Yan, Zeheng Qian, Wangze Ni +6
Cache fusion accelerates generation process of LLMs equipped with RAG through KV caching and selective token recomputation, thereby reducing computational costs and improving effic…
cs.CV2025
V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
Sen Nie, Jie Zhang, Jianxin Yan +2
Adversarial attacks have evolved from simply disrupting predictions on conventional task-specific models to the more complex goal of manipulating image semantics on Large Vision-La…
cs.CL2025
ContextCache: Context-Aware Semantic Cache for Multi-Turn Queries in Large Language Models
Jianxin Yan, Wangze Ni, Lei Chen +4
Semantic caching significantly reduces computational costs and improves efficiency by storing and reusing large language model (LLM) responses. However, existing systems rely prima…