10 papers
ATLAS: Agentic Test-time Learning-to-Allocate Scaling
Peijia Qin, Qi Cao, Pengtao Xie
Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinemen…
AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models
Ruiyi Zhang, Peijia Qin, Qi Cao +2
AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing them remains heavily manual,…
LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling
Qi Cao, Yufan Wang, Peijia Qin +2
Large language models (LLMs) often expose useful signals of self-monitoring: before solving a problem, they can estimate whether they are likely to succeed, and after solving it, t…
BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models
Xin Gao, Ruiyi Zhang, Meixi Du +2
Despite the success of large language models (LLMs) on general-purpose tasks, their performance in highly specialized domains such as biomedicine remains unsatisfactory. A key limi…
AIBuildAI: An AI Agent for Automatically Building AI Models
Ruiyi Zhang, Peijia Qin, Qi Cao +2
AI models underpin modern intelligent systems, driving advances across science, medicine, finance, and technology. Yet developing high-performing AI models remains a labor-intensiv…
Models Under SCOPE: Scalable and Controllable Routing via Pre-hoc Reasoning
Qi Cao, Shuhao Zhang, Ruizhe Zhou +3
Model routing chooses which language model to use for each query. By sending easy queries to cheaper models and hard queries to stronger ones, it can significantly reduce inference…