Showing 2026Show all
2 papers · 1 filter
cs.DC2026
Scaled Block Vecchia Approximation for High-Dimensional Gaussian Process Emulation on GPUs
Qilong Pan, Sameh Abdulah, Mustafa Abduljabbar +8
Emulating computationally intensive scientific simulations is crucial for enabling uncertainty quantification, optimization, and informed decision-making at scale. Gaussian Process…
cs.LG2026
RAP: KV-Cache Compression via RoPE-Aligned Pruning
Jihao Xin, Tian Lyu, David Keyes +2
Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is a direct way to shrink it: dropp…