papers
Publications (2)
cs.LG2026
SCOPE: Signal-Calibrated On-Policy Distillation Enhancement with Dual-Path Adaptive Weighting
Binbin Zheng, Xing Ma, Yiheng Liang +6
On-policy reinforcement learning has become the dominant paradigm for reasoning alignment in large language models, yet its sparse, outcome-level rewards make token-level credit as…
cs.LG2025
PlantBiMoE: A Bidirectional Foundation Model with SparseMoE for Plant Genomes
Kepeng Lin, Qizhe Zhang, Rui Wang +2
Understanding the underlying linguistic rules of plant genomes remains a fundamental challenge in computational biology. Recent advances including AgroNT and PDLLMs have made notab…