2 papers
cs.OS2026
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
Shi Qiu, Yifan Hu, Xintao Wang +6
LLM serving relies on prefix caching to improve inference performance. As growing contexts push key-value (KV) cache footprint far beyond GPU HBM and CPU DRAM capacity, KV cache is…
cs.LG2023
PMNN:Physical Model-driven Neural Network for solving time-fractional differential equations
Zhiying Ma, Jie Hou, Wenhao Zhu +2
In this paper, an innovative Physical Model-driven Neural Network (PMNN) method is proposed to solve time-fractional differential equations. It establishes a temporal iteration sch…