Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
Yuxuan Jiang, Runchao Li, Shubhashis Roy Dipta +2
While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives reasoning gains, an analogous…
cs.CL2026
VeriGraph: Towards Verifiable Data-Analytic Agents
Jiajie Jin, Zhao Yang, Wenle Liao +5
LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes the…