4 papers
Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards
Haoyu Zhang, Yi Feng, Shibo Zheng +6
When an encoded attack passes a content guard, the guard either never represented the payload's harmful content or represented it and failed to act. End-to-end attack success rate…
An Interactive Agent for Requirement-Driven Candidate Sourcing
Yuanpeng He, Fangjing Li, Xiangyu Ru +10
Finding people from a natural-language description (``ML engineers transitioning to research roles in biotech'') is increasingly delegated to LLM agents and framed as information r…
The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
Haoyu Zhang, Yi Feng, Shibo Zheng +6
Vision-Language Model (VLM) safety is expected to depend on what a request asks for. We show that safety-aligned VLMs also key refusal on a property of a request's form: whether an…
0%, 45%, or 99%: A Guardrail's Own Share of the Refusals It Is Credited With
Haoyu Zhang, Xiangchen Guan, Xiao Luo +7
A defended pipeline's refusals have two producers: the guardrail bolted in front of the model, and the model's own alignment. Recovering the split costs nothing, because a guard bl…