collaborators

7 papers

cs.CV2026

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

Yao Huang, Yitong Sun, Huanran Chen +8

Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or…

cs.AI2026

Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

Jennifer Meng Lu, Ruochen Zhang, Isabelle Lee +3

Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design problems to be difficult for humans or c…

cs.CL2025

DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios

Yao Huang, Yitong Sun, Yichi Zhang +3

Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deceptive behav…

cs.CV2025

The Path to Reconciling Quality and Safety in Text-to-Image Generation: Dataset, Method, and Evaluation

Shouwei Ruan, Zhenyu Wu, Yao Huang +5

Content safety is a fundamental challenge for text-to-image (T2I) models, yet prevailing methods enforce a debilitating trade-off between safety and generation quality. We argue th…

cs.CV2025

NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

Yitong Sun, Yao Huang, Ruochen Zhang +4

Despite the impressive generative capabilities of text-to-image (T2I) diffusion models, they remain vulnerable to generating inappropriate content, especially when confronted with…

cs.CR2025

Never compromise with vulnerabilities: a comprehensive survey on AI governance

Yuchu Jiang, Jian Zhao, Yuchen Yuan +64

The rapid advancement of AI has expanded its capabilities across domains, yet introduced critical technical vulnerabilities, such as algorithmic bias and adversarial sensitivity, t…