From the 1 of 12 linked papers with an AI index.
12 papers
EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
Jun Nie, Yonggang Zhang, Qianshu Cai +3
The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from…
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
Jun Nie, Zhiqin Yang, Zhenheng Tang +4
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluatio…
Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment
Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang +16
Zero2Skill is a robot learning system that autonomously collects, verifies, and resets manipulation data while using a large language model to store and reuse human corrections, dr…
A Control Theory of Predictability in Latent World Models
Hanzhe You, Yonggang Zhang, Maohao Ran +6
Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward. Current…
TTHE: Test-Time Harness Evolution
Jun Nie, Yonggang Zhang, Jun Song +5
The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies interm…
Scaling Multi-Hop Training Data via Graph-Constrained Path Selection
Pengyu Chen, Yonggang Zhang, Mingming Chen +3
Endowing large language models with compositional reasoning over specialized documents requires multi-hop training data at scale, where such data rarely exists outside of curated b…