9 papers
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
Rushi Qiang, Changhao Li, Haotian Sun +3
Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment inte…
Revisiting DAgger in the Era of LLM-Agents
Changhao Li, Rushi Qiang, Jiawei Huang +4
Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole trajectory. Existing recipes…
Evolutionary Task Discovery: Advancing Reasoning Frontiers via Skill Composition and Complexity Scaling
Liqin Ye, Yanbin Yin, Michael Galarnyk +3
The reasoning frontier of Large Language Models (LLMs) has advanced significantly through modern post-training paradigms (e.g., Reinforcement Learning from Verifiable Rewards (RLVR…
Exploration-Driven Optimization for Test-Time Large Language Model Reasoning
Changhao Li, Yuchen Zhuang, Chenxiao Gao +4
Post-training techniques combined with inference-time scaling significantly enhance the reasoning and alignment capabilities of large language models (LLMs). However, a fundamental…
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
Zhengding Hu, Mingge Lu, Zhen Wang +8
LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottleneck. We identify that this…
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
Rushi Qiang, Yuchen Zhuang, Anikait Singh +4
While Language Models (LMs) have made significant progress in automating machine learning engineering (MLE), the acquisition of high-quality MLE training data is significantly cons…