activity
20242026
collaborators

14 papers

cs.CL2026

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing

Kexin Chen, Yi Liu, Haonan Zhang +3

As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception monitors either score visible transcripts or…

cs.SE2026

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

Qinyan Zhou, Peixin Zhang, Jun Sun +2

While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that…

cs.LG2026

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

Haonan Zhang, Dongxia Wang, Yi Liu +2

Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector steering methods adjust the mag…

cs.CV2026

FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models

Xinyuan An, Tao Luo, Gengyun Peng +3

Vision-Language-Action (VLA) models are emerging as a cornerstone for robotics, with flow-matching policies like showing great promise in generating smooth, continuous actio…

cs.LG2025

RP-CATE: Recurrent Perceptron-based Channel Attention Transformer Encoder for Industrial Hybrid Modeling

Haoran Yang, Yinan Zhang, Wenjie Zhang +5

Nowadays, industrial hybrid modeling which integrates both mechanistic modeling and machine learning-based modeling techniques has attracted increasing interest from scholars due t…

cs.SE2025

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal

Haonan Zhang, Dongxia Wang, Yi Liu +5

Large Language Models (LLMs) increasingly exhibit over-refusal - erroneously rejecting benign queries due to overly conservative safety measures - a critical functional flaw that u…