2 papers
cs.NI2026
HW-Router: Hardware-Aware Routing for Scalable Multi-LLM Serving
Ahasan Kabir, Jiaqi Xue, Mengxin Zheng +1
Modern large language model (LLM) serving platforms deploy multiple models across different GPUs, requiring routers to direct incoming queries to appropriate LLMs. However, existin…
cs.AI2026
INFRAMIND: Infrastructure-Aware Multi-Agent Orchestration
Ahasan Kabir, Jiaqi Xue, Mengxin Zheng +1
Existing multi-agent LLM orchestration methods, ranging from brute-force ensembles to learned routers, select models and topologies based on task and model features. However, these…