2 papers
cs.AI2026
SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing
Jiacheng Zhang, Haoyu He, Sen Zhang +5
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies, rather than relying on prede…
cs.SE2026
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
Bang Xie, Senjian Zhang, Zhiyuan Peng +3
Large language models have transformed code generation, enabling unprecedented automation in software development. As mobile ecosystems evolve, HarmonyOS has emerged as a critical…