papers
Publications (3)
cs.LG2026
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
Zhiqi Yu, Zhangquan Chen, Mengting Liu +2
Reinforcement Learning with Verifiable Rewards (RLVR), particularly GRPO, has become the standard for eliciting LLM reasoning. However, its efficiency in exploration and difficulty…
cs.LG2026
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
Zhiqi Yu, Xingping Liu, Haobin Mao +4
Grading in large undergraduate STEM courses often yields minimal feedback due to heavy instructional workloads. We present a large-scale empirical study of AI grading on real, hand…
cs.LG2023
A Comprehensive Survey on Source-free Domain Adaptation
Zhiqi Yu, Jingjing Li, Zhekai Du +2
Over the past decade, domain adaptation has become a widely studied branch of transfer learning that aims to improve performance on target domains by leveraging knowledge from the…