1 citations · 1 across the 4 of their papers we have counts for
6 papers
The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems
Yihao Zhang, Kai Wang, Jiangrong Wu +7
Large Language Models (LLMs) face prominent security risks from jailbreaking, a practice that manipulates models to bypass built-in security constraints and generate unethical or u…
Improving Safety Alignment via Balanced Direct Preference Optimization
Shiji Zhao, Mengyang Wang, Shukun Xiong +7
With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from…
Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
Chenxi Liu, Junjie Liang, Yuqi Jia +4
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an effective approach for improving the reasoning abilities of large language models (LLMs). The Group Relative…
AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
Jinghang Shi, Xiaoyu Tang, Yang Huang +4
Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus…
Blackbox Dataset Inference for LLM
Ruikai Zhou, Kang Yang, Xun Chen +3
Today, the training of large language models (LLMs) can involve personally identifiable information and copyrighted material, incurring dataset misuse. To mitigate the problem of d…
Alleviating the Fear of Losing Alignment in LLM Fine-tuning
Kang Yang, Guanhong Tao, Xun Chen +1
Large language models (LLMs) have demonstrated revolutionary capabilities in understanding complex contexts and performing a wide range of tasks. However, LLMs can also answer ques…