works on

From the 2 of 13 linked papers with an AI index.

collaborators

13 papers

cs.AI2026

Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis

Wanxu Cai, Zhengyu Chen, Huaisheng Zhu +3

The paper introduces a time‑truncation harness that limits temporal information during data synthesis, enabling large language models to perform more effective temporal search and…

cs.AI2026

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

Zhengyu Chen, Teng Xiao, Huaisheng Zhu +3

Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, eval…

cs.CL2026

ToFu: A White-Box, Token-Efficient Agent Harness for Researchers

Junhao Ruan, Yuan Ge, Bei Li +7

ToFu is an open‑source, white‑box agentic harness that lets researchers automate codebase reading, file editing, command execution, and tool integration with high token efficiency…

cs.CL2026

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

Ziran Li, Qiang Wang, Zhengyu Chen +4

Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: co…

cs.AI2026

Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

Tianyuan Shi, Canbin Huang, Bei Li +4

Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather th…

cs.LG2026

Teacher-Guided Policy Optimization for On-Policy Reasoning Distillation under Large Policy Divergence

Xinyu Liu, Kechen Jiao, Chunyang Xiao +10

On-policy distillation (OPD) has become a promising paradigm for reasoning-oriented post-training of large language models (LLMs), especially when combined with reinforcement learn…