2 papers
cs.CL2026
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Meijia Chen, Hao Li, Zheng Lu +12
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces…
cs.LG2026
Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching
Alaa Khamis, Alaa Maalouf
Test-time finetuning (TTFT) is a rapidly evolving paradigm that adapts a language model to each prompt by retrieving related sequences, updating the model on them, and then evaluat…