activity
20242026
collaborators

9 papers

cs.CL2026

EvoRoute: Experience-Driven Self-Routing LLM Agent Systems

Guibin Zhang, Haiyang Yu, Kaiming Yang +4

Complex agentic AI systems, powered by a coordinated ensemble of Large Language Models (LLMs), tool and memory modules, have demonstrated remarkable capabilities on intricate, mult…

cs.CL2025

Selective Weak-to-Strong Generalization

Hao Lang, Fei Huang, Yongbin Li

Future superhuman models will surpass the ability of humans and humans will only be able to \textit{weakly} supervise superhuman models. To alleviate the issue of lacking high-qual…

cs.CL2025

EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models

Tao Zou, Xinghua Zhang, Haiyang Yu +3

With the development and widespread application of large language models (LLMs), the new paradigm of "Model as Product" is rapidly evolving, and demands higher capabilities to addr…

cs.AI2025

Socratic-PRMBench: Benchmarking Process Reward Models with Systematic Reasoning Patterns

Xiang Li, Haiyang Yu, Xinghua Zhang +6

Process Reward Models (PRMs) are crucial in complex reasoning and problem-solving tasks (e.g., LLM agents with long-horizon decision-making) by verifying the correctness of each in…

cs.AI2025

DeepSolution: Boosting Complex Engineering Solution Design via Tree-based Exploration and Bi-point Thinking

Zhuoqun Li, Haiyang Yu, Xuanang Chen +6

Designing solutions for complex engineering challenges is crucial in human production activities. However, previous research in the retrieval-augmented generation (RAG) field has n…

cs.CL2025

Debate Helps Weak-to-Strong Generalization

Hao Lang, Fei Huang, Yongbin Li

Common methods for aligning already-capable models with desired behavior rely on the ability of humans to provide supervision. However, future superhuman models will surpass the ca…