2 papers
stat.ME2025
Regularised Canonical Correlation Analysis: graphical lasso, biplots and beyond
Lennie Wells, Kumar Thurimella, Sergio Bacallado
Recent developments in regularized Canonical Correlation Analysis (CCA) promise powerful methods for high-dimensional, multiview data analysis. However, justifying the structural a…
cs.CL2025
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
Jason R Brown, Lennie Wells, Edward James Young +1
Proximal Policy Optimisation (PPO) is an established and effective policy gradient algorithm used for Language Model Reinforcement Learning from Human Feedback (LM-RLHF). PPO perfo…