3 papers
cs.LG2026
Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
Hu Wang, Congbo Ma, Ian Reid +1
The advantage function is a central concept in RL that helps reduce variance in policy gradient estimates. For language modeling, Group Relative Policy Optimization (GRPO) was prop…
cs.LG2025
Rethinking Weight-Averaged Model-merging
Hu Wang, Congbo Ma, Ibrahim Almakky +3
Model merging, particularly through weight averaging, has shown surprising effectiveness in saving computations and improving model performance without any additional training. How…
cs.CV2025
ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang +3
The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised l…