3 papers
cs.AI2026
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
Shihao Weng, Yang Feng, Xiaofei Xie
LLM-as-a-Judge pipelines have become the de facto evaluator for agent safety, yet existing benchmarks treat their verdicts as ground-truth proxies without checking whether the verd…
cs.SE2025
Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
Wei Ma, Yixiao Yang, Jingquan Ge +2
Code generation models are widely used in software development, yet their sensitivity to prompt phrasing remains under-examined. Identical requirements expressed with different emo…
cs.CY2024
Navigating Governance Paradigms: A Cross-Regional Comparative Study of Generative AI Governance Processes & Principles
Jose Luna, Ivan Tan, Xiaofei Xie +1
As Generative Artificial Intelligence (GenAI) technologies evolve at an unprecedented rate, global governance approaches struggle to keep pace with the technology, highlighting a c…