3 papers
cs.LG2026
GAPO: Robust Advantage Estimation for Real-World Code LLMs
Jianqing Zhang, Zhezheng Hao, Wei Xia +7
Reinforcement learning (RL) is widely used for post-training large language models (LLMs) in code editing, where group-relative methods, such as GRPO, are popular due to their crit…
cs.SE2026
AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
Jianqing Zhang, Wei Xia, Hande Dong +2
LLMs' code generation capabilities have yielded substantial improvements in the effectiveness of programming tasks. However, LLM-generated code still suffers from compilation and r…
cs.CV2025
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
Wei Guan, Jun Lan, Jian Cao +3
Industrial anomaly detection (IAD) plays a crucial role in maintaining the safety and reliability of manufacturing systems. While multimodal large language models (MLLMs) show stro…