3 papers
cs.LG2026
ScDiVa: Masked Discrete Diffusion for Joint Modeling of Single-Cell Identity and Expression
Mingxuan Wang, Cheng Chen, Gaoyang Jiang +4
Single-cell RNA-seq profiles are high-dimensional, sparse, and unordered, causing autoregressive generation to impose an artificial ordering bias and suffer from error accumulation…
cs.LG2025
Curiosity Meets Cooperation: A Game-Theoretic Approach to Long-Tail Multi-Label Learning
Canran Xiao, Chuangxin Zhao, Zong Ke +1
Long-tail imbalance is endemic to multi-label learning: a few head labels dominate the gradient signal, while the many rare labels that matter in practice are silently ignored. We…
cs.CL2025
-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
Yining Wang, Jinman Zhao, Chuangxin Zhao +3
Reinforcement Learning with Human Feedback (RLHF) has been the dominant approach for improving the reasoning capabilities of Large Language Models (LLMs). Recently, Reinforcement L…