2 papers
cs.LG2026
CacheClip: Accelerating RAG with Effective KV Cache Reuse
Bin Yang, Qiuyu Leng, Jun Zeng +1
Retrieval-Augmented Generation (RAG) systems suffer from severe time-to-first-token (TTFT) bottlenecks due to long input sequences. Existing KV cache reuse methods face a fundament…
cs.CE2026
A Generalizable Framework for Building Executable Domain-Specific LLMs under Data Scarcity: Demonstration on Semiconductor TCAD Simulation
Di Wang, Zhenhua Wu, Yu Liu +2
Scientific and engineering verticals often suffer from data scarcity and strict executability requirements: models must generate not only fluent text, but also syntactically valid,…