6 papers
Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis
Wanxu Cai, Zhengyu Chen, Huaisheng Zhu +3
The paper introduces a time‑truncation harness that limits temporal information during data synthesis, enabling large language models to perform more effective temporal search and…
ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses
Yusheng Zheng, Tianyuan Wu, Quanzhi Fu +6
AI agents increasingly run in production through harnesses, the software around the LLM, including an engine that enforces safety and effectiveness policies, e.g., 'run tests befor…
Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests
Jin-Guo Liu, Shang-Qi Lu, Xin-Ran Shi +2
This paper is an experience report on a 13-week Test-Driven, AI-Assisted (TDAA) redesign of DSAA 3071, Theory of Computation, an upper-level course at the Hong Kong University of S…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Reinforcement Learning for Tool-Integrated Interleaved Thinking towards Cross-Domain Generalization
Zhengyu Chen, Jinluan Yang, Teng Xiao +6
Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in reasoning and tool utilization. However, the generalization of tool-augmented reinforce…
From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
Zhengyu Chen, Yudong Wang, Teng Xiao +5
Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors…