activity
20242026
collaborators

7 papers

cs.CV2026

PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails

Mingyang Song, Luxin Xu, Haoyu Sun +3

Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: t…

cs.LG2026

CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes

Zhijian Zhou, Zesheng Ye, Zhaorun Chen +2

Can we trust evaluation scores to capture an LLM's true real-world performance? Certifiable evaluation answers this question by providing guarantee for LLM evaluation. In particula…

cs.CR2026

Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset

Mintong Kang, Zhaorun Chen, Chejian Xu +6

As LLMs become widespread across diverse applications, concerns about the security and safety of LLM interactions have intensified. Numerous guardrail models and benchmarks have be…

cs.LG2025

ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning

Zhaorun Chen, Mintong Kang, Bo Li

Autonomous agents powered by foundation models have seen widespread adoption across various real-world applications. However, they remain highly vulnerable to malicious instruction…

cs.CV2025

SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability

Peiyang Xu, Minzhou Pan, Zhaorun Chen +3

With the rapid proliferation of digital media, the need for efficient and transparent safeguards against unsafe content is more critical than ever. Traditional image guardrail mode…

stat.ML2025

Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

Ganghua Wang, Zhaorun Chen, Bo Li +1

As foundation models continue to scale, the size of trained models grows exponentially, presenting significant challenges for their evaluation. Current evaluation practices involve…