4 papers
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
Rushi Qiang, Changhao Li, Haotian Sun +3
Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment inte…
DREAM: Deep Research Evaluation with Agentic Metrics
Elad Ben Avraham, Changhao Li, Ron Dorfman +8
Deep Research Agents generate analyst-grade reports, yet evaluating them remains challenging due to the absence of a single ground truth and the multidimensional nature of research…
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
Changhao Li, Yuchen Zhuang, Rushi Qiang +4
Despite the impressive generative abilities of black-box large language models (LLMs), their inherent opacity hinders further advancements in capabilities such as reasoning, planni…
MLE-Dojo: Interactive Environments for Empowering LLM Agents in Machine Learning Engineering
Rushi Qiang, Yuchen Zhuang, Yinghao Li +8
We introduce MLE-Dojo, a Gym-style framework for systematically reinforcement learning, evaluating, and improving autonomous large language model (LLM) agents in iterative machine…