2 papers
cs.LG2024
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
Weizhuo Li, Zhigang Wang, Yu Gu +1
Recently the generative Large Language Model (LLM) has achieved remarkable success in numerous applications. Notably its inference generates output tokens one-by-one, leading to ma…
cs.LG2024
Hierarchical Gradient-Based Genetic Sampling for Accurate Prediction of Biological Oscillations
Heng Rao, Yu Gu, Jason Zipeng Zhang +3
Biological oscillations are periodic changes in various signaling processes crucial for the proper functioning of living organisms. These oscillations are modeled by ordinary diffe…