activity
20242026
most citedMellivora Capensis: A Backdoor-Free Training Framework on the Poisoned Dataset without Auxiliary Data

1 citations · 2 across the 24 of their papers we have counts for

collaborators
Showing cs.CRShow all

20 papers · 1 filter

cs.CR2026

A Finger on the Scale: Covert Policy Steering through Agentic Skills

Jiarui Li, Jiahao Chen, Chunyi Zhou +5

Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also act as externalized behavioral…

cs.CR2026

The Shape of Ownership: Verifying LLM Provenance through Semantic Structures

Zhongrui Sun, Jiahao Chen, Oubo Ma +4

As large language models (LLMs) are increasingly redistributed, adapted, and served behind opaque APIs, model ownership can no longer be established reliably by inspecting model in…

cs.CR2026

Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds

Jiahao Chen, Rui Yin, Xinfeng Li +6

Large Language Models (LLMs) have been integrated into complex ecosystems (e.g., Code Agents), while Indirect Prompt Injection (IPI) attacks have emerged as critical barriers to th…

cs.CR2026

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails

Shuo Shi, Rui Yin, Naen Xu +7

Multimodal guard models have emerged as critical safety components for screening content in vision-language systems. While adversarial research has extensively studied jailbreaking…

cs.CR2026

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

Zhou Feng, Jiahao Chen, Chunyi Zhou +6

Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persisten…

cs.CR2026

Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks

Yu Yan, Jiahao Chen, Siqi Lu +6

Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. Howe…