3 papers
stat.ML2026
Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits
Kaifei Wang, Yinyu Ye, Han Zhong
Multi-armed bandit algorithms are evaluated by regret, yet comparable regret can coexist with different allocations across independent runs. We study the trade-off between worst-ca…
cs.LG2025
Reasoning in Diffusion Large Language Models is Concentrated in Dynamic Confusion Zones
Ranfei Chen, Ming Chen, Kaifei Wang
Diffusion Large Language Models (dLLMs) are rapidly emerging alongside autoregressive models as a powerful paradigm for complex reasoning, with reinforcement learning increasingly…
cs.LG2025
pUniFind: a unified large pre-trained deep learning model pushing the limit of mass spectra interpretation
Jiale Zhao, Pengzhi Mao, Kaifei Wang +12
Deep learning has advanced mass spectrometry data interpretation, yet most models remain feature extractors rather than unified scoring frameworks. We present pUniFind, the first l…