Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PHF: Privileged Hidden Flow for On-Policy Self-Distillation
Yuhan Li, Mingxu Zhang, Dazhong Shen +1
On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions. Ex…
cs.AI2025
GraphArena: Evaluating and Exploring Large Language Models on Graph Computation
Jianheng Tang, Qifan Zhang, Yuhan Li +2
The ``arms race'' of Large Language Models (LLMs) demands new benchmarks to examine their progresses. In this paper, we introduce GraphArena, a benchmarking tool designed to evalua…