2 papers
cs.AI2026
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
Shihao Weng, Yang Feng, Xiaofei Xie
LLM-as-a-Judge pipelines have become the de facto evaluator for agent safety, yet existing benchmarks treat their verdicts as ground-truth proxies without checking whether the verd…
cs.SE2025
Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
Wei Ma, Yixiao Yang, Jingquan Ge +2
Code generation models are widely used in software development, yet their sensitivity to prompt phrasing remains under-examined. Identical requirements expressed with different emo…