2 papers
cs.CR2026
Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions
Junjie Xiong, Zhengyuan Jiang, Xiaoran Xu +7
Large Language Models (LLMs) have emerged as powerful tools that impact information integrity on social media platforms. This comprehensive review examines the dual role of LLMs in…
cs.CR2026
Prompt Overflow: What the Guardrail Inspects Is Not What the Model Infers
Yuanbo Zhou, Changjia Zhu, Junyu Wang +5
Guardrail models (a.k.a. safety checkers) are widely deployed to screen user inputs before they reach large language models (LLMs), serving as a primary defense against prompt inje…