1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2025★ 1 cited
Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
Yingshui Tan, Yilei Jiang, Yanshi Li +6
Fine-tuning large language models (LLMs) based on human preferences, commonly achieved through reinforcement learning from human feedback (RLHF), has been effective in improving th…
cs.CL2024
RapGuard: Safeguarding Multimodal Large Language Models via Rationale-aware Defensive Prompting
Yilei Jiang, Yingshui Tan, Xiangyu Yue
While Multimodal Large Language Models (MLLMs) have made remarkable progress in vision-language reasoning, they are also more susceptible to producing harmful content compared to m…