7 papers
LatentRevise: Learning from Zero-Hit Reasoning
Yiqiu Guo, Xueting Han, Qi Jia +2
Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by hard prompts on which correct trajectories have low probability, so sampling misses them within a practical…
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
Liang Chen, Xueting Han, Li Shen +2
Supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) are two widely used post-training paradigms for improving the reasoning ability of large lang…
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
Liang Chen, Xueting Han, Qizhou Wang +4
Balancing exploration and exploitation remains a central challenge in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs). Current RLVR methods o…
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
Yiyang Jin, Kunzhao Xu, Hang Li +4
Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. However, existing methods rely solely on outcome rewards, wi…
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
Liang Chen, Xueting Han, Li Shen +2
Harmful fine-tuning (HFT), performed directly on open-source LLMs or through Fine-tuning-as-a-Service, breaks safety alignment and poses significant threats. Existing methods aim t…
BSM: Small but Powerful Biological Sequence Model for Genes and Proteins
Weixi Xiang, Xueting Han, Xiujuan Chai +1
Modeling biological sequences such as DNA, RNA, and proteins is crucial for understanding complex processes like gene regulation and protein synthesis. However, most current models…