activity
20242026
collaborators

8 papers

cs.CL2026

Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation

Tong Li, Shu Yang, Junchao Wu +6

We present a comprehensive evaluation framework for assessing Large Language Models' (LLMs) capabilities in suicide prevention, focusing on two critical aspects: the Identification…

cs.LG2025

PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning

Mengdi Li, Guanqiao Chen, Xufeng Zhao +3

Reward models (RMs), which are central to existing post-training methods, aim to align LLM outputs with human values by providing feedback signals during fine-tuning. However, exis…

cs.LG2025

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models

Liangyu Wang, Huanyi Xie, Xinhai Wang +3

Group-based reinforcement learning algorithms such as Group Reward Policy Optimization (GRPO) have proven effective for fine-tuning large language models (LLMs) with human feedback…

cs.CL2025

Understanding the Repeat Curse in Large Language Models from a Feature Perspective

Junchi Yao, Shu Yang, Jianhua Xu +3

Large language models (LLMs) have made remarkable progress in various domains, yet they often suffer from repetitive text generation, a phenomenon we refer to as the "Repeat Curse"…

cs.CL2025

Fraud-R1 : A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements

Shu Yang, Shenzhe Zhu, Zeyu Wu +7

We introduce Fraud-R1, a benchmark designed to evaluate LLMs' ability to defend against internet fraud and phishing in dynamic, real-world scenarios. Fraud-R1 comprises 8,564 fraud…

cs.LG2025

Towards User-level Private Reinforcement Learning with Human Feedback

Jiaming Zhang, Mingxi Lei, Meng Ding +5

Reinforcement Learning with Human Feedback (RLHF) has emerged as an influential technique, enabling the alignment of large language models (LLMs) with human preferences. Despite th…