7 papers
ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision-Language Models
Zhihao Dou, Qinjian Zhao, Zhiqiang Gao +1
Vision--Language Models (VLMs) are increasingly deployed in safety-critical applications, yet remain vulnerable to backdoor attacks. Existing methods primarily manipulate final out…
The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models
Alif Al Hasan, Sumon Biswas
Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both failure modes across 21 open-weig…
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
Alif Al Hasan, Sumon Biswas
Autonomous coding agents built on large language models (LLMs) are rapidly being integrated into development workflows, yet their operational safety properties remain poorly unders…
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
Zhihao Dou, Qinjian Zhao, Zhongwei Wan +10
Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global plannin…
CoRe-Code: Collaborative Reinforcement Learning for Code Generation
Zhihao Dou, Qinjian Zhao, Zhongwei Wan +2
Large language models (LLMs) have achieved strong performance in code generation, but most methods rely on autoregressive decoding without global planning, often leading to locally…
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
Sina Salimian, Gias Uddin, Sumon Biswas +1
The widespread deployment of Large Language Models (LLMs) has intensified concerns about subtle social biases embedded in their outputs. Existing guardrails often fail when faced w…