8 papers
One Run Is Not an Idea: The Implementation Lottery in Automated Research
Jingjie Ning, Shanshan Zhong, Xiaochuan Li +2
The paper studies how automated research systems can draw misleading conclusions when they rely on a single implementation of an idea, introducing the concept of an "implementation…
ACM: Agentic Context Management for Long Horizon Tasks
Xiaochuan Li, Ryan Ming, Meng Chu +3
Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existing context compression methods inevitably…
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
Jingjie Ning, Xiaochuan Li, Shanshan Zhong +2
Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged by its terminal pipeline. A terminal s…
Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements
Jingjie Ning, Xiaochuan Li, Ji Zeng +2
Closed-loop Auto Research extends automated machine learning from fixed-dataset fitting to changing the research workflow, with language-model agents editing representations and mo…
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
Jingjie Ning, Xiaochuan Li, Ji Zeng +2
We study auto research as a closed empirical loop driven by external measurement. Each submitted trial carries a hypothesis, an executable code edit, an evaluator-owned outcome, an…
Benchmark Test-Time Scaling of General LLM Agents
Xiaochuan Li, Ryan Ming, Pranav Setlur +6
LLM agents are increasingly expected to function as general-purpose systems capable of resolving open-ended user requests. While existing benchmarks focus on domain-aware environme…