5 papers
RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing
Guannan Lai, Haoran Hu, Han-Jia Ye
We present RouteJudge, an online pairwise preference evaluation framework for LLM routing systems, with a public platform available at https://routejudge.cn. Different from model-l…
From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing
Guannan Lai, Haoran Hu, Long Chen +2
Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers. However, because LLM generation is inherently stocha…
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
Qiyao Wang, Haoran Hu, Longze Chen +4
With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code sy…
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
Yinghao Zhu, Ziyi He, Haoran Hu +6
The rapid advancement of Large Language Models (LLMs) has stimulated interest in multi-agent collaboration for addressing complex medical tasks. However, the practical advantages o…
HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
Yinghao Zhu, Yifan Qi, Zixiang Wang +8
The rapid proliferation of scientific knowledge presents a grand challenge: transforming this vast repository of information into an active engine for discovery, especially in high…