Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Attention Is All You Need for KV Cache in Diffusion LLMs
Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
This work studies how to adaptively recompute key-value (KV) caches for diffusion large language models (DLMs) to maximize prediction accuracy while minimizing decoding latency. Pr…
cs.CL2025
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
Sondos Mahmoud Bsharat, Mukul Ranjan, Aidar Myrzakhan +6
Rapid advancements in large language models (LLMs) have increased interest in deploying them on mobile devices for on-device AI applications. Mobile users interact differently with…