3 papers
cs.CL2026
Are LLMs Positionally Consistent Ordinal Classifiers? A Systematic Evaluation
Yu Wang, Zhe Zhou, Menglin Liu +1
Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systemati…
cs.CY2026
What a Model Refuses, a State Fears: How Authoritarian Information Control Reproduces in Language-Model Guardrails
Menglin Liu, Yao Yu, Tong Wu +2
As large language models become the front door to political information, what they refuse to discuss becomes a new instrument of information control. We argue that a model's guardr…
cs.CR2026
JailbreakOPT: Tool-Assisted Iterative Jailbreak Prompt Optimization
Ge Shi, Jun Yin, Donglin Xie +3
Jailbreak attacks expose persistent safety weaknesses in large language models (LLMs), but existing stateless single-turn methods face a trade-off: hand-crafted prompts are express…