11 papers
AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
Zhangchen Xu, Junda Chen, Yue Huang +16
Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifac…
Potential-Guided Flow Matching for Vision-Language-Action Policy Improvement
Yunpeng Mei, Jiakai He, Hongjie Cao +12
Large vision-language-action (VLA) policies are increasingly trained as conditional generative models over action chunks. Yet deployment produces mixed-quality experience-successfu…
MAD: Mapping-Aware World Models for Agile Quadrotor Flight
Xinhong Zhang, Runqing Wang, Yunfan Ren +6
Agile quadrotor flight in cluttered scenes requires more than a reactive mapping from a depth image to a control command: the vehicle must remember which regions have been observed…
Converted, Not Equivalent: Benchmarking Codebase Conversion via Observational Equivalence
Linxin Song, Jiefeng Chen, Yue Huang +5
Coding agents increasingly act as codebase-scale collaborators that can assist with codebase conversion, but this progress has exposed a critical weakness: agents often over-trust…
MARS: Modular Agent with Reflective Search for Automated AI Research
Jiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam +3
A critical bottleneck in automating AI research is the execution of complex machine learning engineering (MLE) tasks. MLE differs from general software engineering due to computati…
DS-STAR: Data Science Agent for Solving Diverse Tasks across Heterogeneous Formats and Open-Ended Queries
Jaehyun Nam, Jinsung Yoon, Jiefeng Chen +3
While large language models (LLMs) have shown promise in automating data science, existing agents often struggle with the complexity of real-world workflows that require exploring…