2 papers
cs.OS2026
CloakLM: Obfuscating GPU Memory Layout to Mitigate Model Ex-filtration for Serving
Kunal Jain, Seokjin Go, Divya Mahajan
Large foundation models deployed on third-party and shared accelerator infrastructure face a practical risk of model exfiltration that existing defenses do not fully address. In co…
cs.AR2025
Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
Wonung Kim, Yubin Lee, Yoonsung Kim +8
Transformers are the driving force behind today's Large Language Models (LLMs), serving as the foundation for their performance and versatility. Yet, their compute and memory costs…