3 papers
cs.CL2026
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
Guochao Jiang, Jingyi Song, Guofeng Quan +3
Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative Policy Optimization offers an…
cs.CR2025
PortGPT: Towards Automated Backporting Using Large Language Models
Zhaoyang Li, Zheng Yu, Jingyi Song +3
Patch backporting, the process of migrating mainline security patches to older branches, is an essential task in maintaining popular open-source projects (e.g., Linux kernel). Howe…
cs.AI2025
AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
Wenfeng Feng, Chuzhan Hao, Yuewei Zhang +3
Leveraging the autonomous decision-making capabilities of large language models (LLMs) has demonstrated superior performance in reasoning tasks. However, despite the success of ite…