Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
Peng Kuang, Yanli Wang, Xiaoyu Han +3
Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise…
cs.CL2026
Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting
Devang Kulshreshtha, Hang Su, Haibo Jin +2
We introduce \emph{self-jailbreaking}, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which often rely on handcrafted prompts o…
cs.CL2025
Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes
Sullam Jeoung, Yubin Ge, Haohan Wang +1
Examining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs wi…