Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Ascendra: Dynamic Request Prioritization for Efficient LLM Serving
Azam Ikram, Xiang Li, Sameh Elnikety +1
The rapid advancement of Large Language Models (LLMs) has driven the need for more efficient serving strategies. In this context, efficiency refers to the proportion of requests th…
cs.AI2025
Learning to Inference Adaptively for Multimodal Large Language Models
Zhuoyan Xu, Khoi Duc Nguyen, Preeti Mukherjee +4
Multimodal Large Language Models (MLLMs) have shown impressive capabilities in visual reasoning, yet come with substantial computational cost, limiting their deployment in resource…