4 papers
Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
Yu-Lin Tsai, Yu-An Lu, Ci-Yang Tsai +3
Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. Hosted providers can keep these files secret while sell…
Uncertainty-Aware Reward Modeling for Stable RLHF
Licheng Pan, Haocheng Yang, Haoxuan Li +7
Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…
Robust Reward Modeling for Large Language Models via Causal Decomposition
Yunsheng Lu, Zijiang Yang, Licheng Pan +1
Reward models are central to aligning large language models, yet they often overfit to spurious cues such as response length and overly agreeable tone. Most prior work weakens thes…
A Causal Perspective for Enhancing Jailbreak Attack and Defense
Licheng Pan, Yunsheng Lu, Jiexi Liu +5
Uncovering the mechanisms behind "jailbreaks" in large language models (LLMs) is crucial for enhancing their safety and reliability, yet these mechanisms remain poorly understood.…