2 papers
cs.CL2026
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression
Shen Han, Yuyang Wu, Junpu Yu +1
Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited th…
cs.AI2026
CrystalReasoner: Reasoning and RL for Property-Conditioned Crystal Structure Generation
Yuyang Wu, Stefano Falletta, Delia McGrath +1
Generative modeling has emerged as a promising approach for crystal structure discovery. However, existing LLM-based generative models struggle with low-level atomic precision, whi…