most citedBeyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

2 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.AI2026

LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing

Hao Li, Yiqun Zhang, Zhaoyan Guo +9

Large language model (LLM) routing assigns each query to the most suitable model from an ensemble. We introduce LLMRouterBench, a large-scale benchmark and unified framework for LL…

cs.LG2025

ICL-Router: In-Context Learned Model Representations for LLM Routing

Chenxu Wang, Hao Li, Yiqun Zhang +6

Large language models (LLMs) often exhibit complementary strengths. Model routing harnesses these strengths by dynamically directing each query to the most suitable model, given a…

cs.AI2025

PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration

Xin Wang, Zhiyao Cui, Hao Li +10

Vision language model (VLM)-based mobile agents show great potential for assisting users in performing instruction-driven tasks. However, these agents typically struggle with perso…

cs.CL20252 cited

Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

Yiqun Zhang, Hao Li, Jianhao Chen +4

Balancing performance and efficiency is a central challenge in large language model (LLM) advancement. GPT-5 addresses this with test-time routing, dynamically assigning queries to…

cs.AI2025

Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior

Hao Li, Gengrui Zhang, Petter Holme +2

Human decision-making belongs to the foundation of our society and civilization, but we are on the verge of a future where much of it will be delegated to artificial intelligence.…

cs.CL2025

The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants

Yiqun Zhang, Hao Li, Chenxu Wang +11

Proprietary giants are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this p…