2 papers
cs.LG2025
In-depth Analysis on Caching and Pre-fetching in Mixture of Experts Offloading
Shuning Lin, Yifan He, Yitong Chen
In today's landscape, Mixture of Experts (MoE) is a crucial architecture that has been used by many of the most advanced models. One of the major challenges of MoE models is that t…
cs.CL2025
Cash or Comfort? How LLMs Value Your Inconvenience
Mateusz Cedro, Timour Ichmoukhamedov, Sofie Goethals +3
Large Language Models (LLMs) are increasingly proposed as near-autonomous artificial intelligence (AI) agents capable of making everyday decisions on behalf of humans. Although LLM…