7 papers
Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial Nets
Hyunjoo Kim, Sicheng Wu, Agastya Venkatraman +2
Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecified statistical models. One impo…
Generalized Discrete Diffusion with Self-Correction
Linxuan Wang, Ziyi Wang, Yikun Bai +3
Self-correction is an effective technique for maintaining parallel sampling in discrete diffusion models with minimal performance degradation. Prior work has explored self-correcti…
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment
Rajdeep Haldar, Lantao Mei, Guang Lin +2
Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) & unaligned (less-preferred) distributions, yielding…
Task-tailored Pre-processing: Fair Downstream Supervised Learning
Jinwon Sohn, Guang Lin, Qifan Song
Fairness-aware machine learning has recently attracted various communities to mitigate discrimination against certain societal groups in data-driven tasks. For fair supervised lear…
Trusted Multi-view Learning for Long-tailed Classification
Chuanqing Tang, Yifei Shi, Guanghao Lin +2
Class imbalance has been extensively studied in single-view scenarios; however, addressing this challenge in multi-view contexts remains an open problem, with even scarcer research…
LLM Safety Alignment is Divergence Estimation in Disguise
Rajdeep Haldar, Ziyi Wang, Qifan Song +2
We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or…