3 papers
cs.LG2026
A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization
Prasanth YSS, Zhichen Ren, Rasa Hosseinzadeh +6
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization remains prone to collapse. We analyse this instability through…
q-bio.BM2025
Protein Large Language Models: A Comprehensive Survey
Yijia Xiao, Wanjia Zhao, Junkai Zhang +12
Protein-specific large language models (Protein LLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design.…
cs.SI2024
Do We Trust What They Say or What They Do? A Multimodal User Embedding Provides Personalized Explanations
Zhicheng Ren, Zhiping Xiao, Yizhou Sun
With the rapid development of social media, the importance of analyzing social network user data has also been put on the agenda. User representation learning in social media is a…