activity
20242026
most citedTensor Product Attention Is All You Need

4 citations · 6 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Fast Weight Attention for Continual Learning

Yifan Zhang, Steve Ta, Jasper Zhang +8

Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule…

cs.LG2025

Group Representational Position Encoding

Yifan Zhang, Zixiang Chen, Yifeng Liu +6

We present GRAPE (Group Representational Position Encoding), a unified framework for positional encoding based on group actions. GRAPE unifies two families of mechanisms: (i) multi…

cs.LG2025

On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning

Yifan Zhang, Yifeng Liu, Huizhuo Yuan +3

Policy gradient algorithms have been successfully applied to enhance the reasoning capabilities of large language models (LLMs). KL regularization is ubiquitous, yet the design sur…

cs.LG20251 cited

ConfRover: Simultaneous Modeling of Protein Conformation and Dynamics via Autoregression

Yuning Shen, Lihao Wang, Huizhuo Yuan +3

Understanding protein dynamics is critical for elucidating their biological functions. The increasing availability of molecular dynamics (MD) data enables the training of deep gene…

cs.LG2025

RSPO: Regularized Self-Play Alignment of Large Language Models

Xiaohang Tang, Sangwoong Yoon, Seongho Son +3

Self-play-based policy optimization has emerged as an effective approach for fine-tuning large language models (LLMs), formulating preference optimization as a two-player game. How…

cs.LG2024

Towards Simple and Provable Parameter-Free Adaptive Gradient Methods

Yuanzhe Tao, Yifeng Liu, Huizhuo Yuan +3

Optimization algorithms such as AdaGrad and Adam have significantly advanced the training of deep models by dynamically adjusting the learning rate during the optimization process.…