3 papers
cs.AI2026
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
Yifei Wang, Tianlin Li, Xiaohan Zhang +4
Large language models (LLMs) are increasingly deployed under diverse numerical precision configurations, including standard floating-point formats (e.g., bfloat16 and float16) and…
cs.CL2025
SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation
Lai Jiang, Yuekang Li, Xiaohan Zhang +2
Accurate jailbreak evaluation is critical for LLM red team testing and jailbreak research. Mainstream methods rely on binary classification (string matching, toxic text classifiers…
cs.AI2025
MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair
Changqing Li, Tianlin Li, Xiaohan Zhang +2
Large Language Models (LLMs) face persistent and evolving trustworthiness issues, motivating developers to seek automated and flexible repair methods that enable convenient deploym…