2 papers
cs.DC2025
Accelerating LLM Inference with Precomputed Query Storage
Jay H. Park, Youngju Cho, Choungsol Lee +2
Large language model (LLM) inference often suffers from high latency, particularly in resource-constrained environments such as on-device or edge deployments. To address this chall…
cs.LG2025
Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)
Shengzhe Xu, Cho-Ting Lee, Mandar Sharma +3
Synthetic data generation is integral to ML pipelines, e.g., to augment training data, replace sensitive information, and even to power advanced platforms like DeepSeek. While LLMs…