3 papers
cs.CL2026
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
Mingyu Luo, Ming Deng, Zilang Qiu +8
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read…
cs.SD2026
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Wanyi Ning, Wei Zhou, Yingpeng Li +3
Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision ar…
cs.CL2026
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
Y. H. Zhou, Z. M. Ma, Y. J. Zhou +12
SMS fraud is increasingly cross-channel: a message directs the user to a webpage, and the final risk depends on how the SMS claim aligns with the page content and requested user ac…