activity
20232026
most citedSelf-Play Preference Optimization for Language Model Alignment

12 citations · 40 across the 15 of their papers we have counts for

collaborators

15 papers

cs.LG2026

Fast Weight Attention for Continual Learning

Yifan Zhang, Steve Ta, Jasper Zhang +8

Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule…

cs.LG2025

Group Representational Position Encoding

Yifan Zhang, Zixiang Chen, Yifeng Liu +6

We present GRAPE (Group Representational Position Encoding), a unified framework for positional encoding based on group actions. GRAPE unifies two families of mechanisms: (i) multi…

cs.CL2025

Causal Attention with Lookahead Keys

Zhuoqing Song, Peng Sun, Huizhuo Yuan +1

In standard causal attention, each token's query, key, and value (QKV) are static and encode only preceding context. We introduce CAuSal aTtention with Lookahead kEys (CASTLE), an…

cs.LG2025

On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning

Yifan Zhang, Yifeng Liu, Huizhuo Yuan +3

Policy gradient algorithms have been successfully applied to enhance the reasoning capabilities of large language models (LLMs). KL regularization is ubiquitous, yet the design sur…

cs.LG2025★ 1 cited

ConfRover: Simultaneous Modeling of Protein Conformation and Dynamics via Autoregression

Yuning Shen, Lihao Wang, Huizhuo Yuan +3

Understanding protein dynamics is critical for elucidating their biological functions. The increasing availability of molecular dynamics (MD) data enables the training of deep gene…

cs.LG2025

RSPO: Regularized Self-Play Alignment of Large Language Models

Xiaohang Tang, Sangwoong Yoon, Seongho Son +3

Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game. How…