2 papers
cs.NI2026
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
Yasmin Moslem, John D. Kelleher
The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent model selection at inference time. While…
cs.CL2025
Iterative Layer Pruning for Efficient Translation Inference
Yasmin Moslem, Muhammad Hazim Al Farouq, John D. Kelleher
Large language models (LLMs) have transformed many areas of natural language processing, including machine translation. However, efficient deployment of LLMs remains challenging du…