collaborators

18 papers

cs.CR2026

AMRM-Pure: Semantic-Preserving Adversarial Purification

Zhihao Dou, Zhiqiang Gao, Dongfei Cui +6

Adversarial purification is a defense technique that employs generative models to remove adversarial perturbations. Current methods often rely on powerful generators, typically dif…

cs.CR2026

ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision-Language Models

Zhihao Dou, Qinjian Zhao, Zhiqiang Gao +1

Vision--Language Models (VLMs) are increasingly deployed in safety-critical applications, yet remain vulnerable to backdoor attacks. Existing methods primarily manipulate final out…

cs.AI2026

STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning

Qinjian Zhao, Zhihao Dou, Dinggen Zhang +10

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities of large language models. However, existi…

cs.AI2026

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning

Zhihao Dou, Qinjian Zhao, Zhongwei Wan +10

Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global plannin…

cs.AI2026

CoRe-Code: Collaborative Reinforcement Learning for Code Generation

Zhihao Dou, Qinjian Zhao, Zhongwei Wan +2

Large language models (LLMs) have achieved strong performance in code generation, but most methods rely on autoregressive decoding without global planning, often leading to locally…

cs.AI2026

SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills

Yingtie Lei, Zhongwei Wan, Jiankun Zhang +13

Large language model (LLM) agents accumulate rich episodic trajectories while solving real-world tasks, but it remains unclear whether such experience can be distilled into reusabl…