8 papers
EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
Jun Nie, Yonggang Zhang, Qianshu Cai +3
The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from…
Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection
Jun Nie, Yonggang Zhang, Tongliang Liu +3
Robust detection of generated images is critical to counter the misuse of generative models. Existing methods primarily depend on learning from human-annotated training datasets, l…
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
Jun Nie, Zhiqin Yang, Zhenheng Tang +4
Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluatio…
TTHE: Test-Time Harness Evolution
Jun Nie, Yonggang Zhang, Jun Song +5
The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies interm…
PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking
Junnan Nie, Jiayi Li, Jiachen Zhang +5
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an…
What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies
Jiachen Zhang, Junnan Nie, Junyi Lao +4
Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Their frozen representations nev…