Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
Jun Nie, Yonggang Zhang, Qianshu Cai +3
The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from…
cs.LG2026
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
Jun Nie, Zhiqin Yang, Zhenheng Tang +4
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluatio…