collaborators

9 papers

cs.CL2026

Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety

Ping Wu, Haibo Tong, Feifei Zhao +7

Safety tuning can improve harmful refusal, but models may learn surface-form shortcuts: wrapped harmful prompts bypass safety, while similarly wrapped benign prompts are over-refus…

cs.NE2026

Spiking Local Interaction and Adaptive Complementary Fusion for Spiking Transformer

Dongcheng Zhao, Sicheng Shen, Zhenyu Yang +6

Spiking Transformers model token interactions primarily through spiking self-attention (SSA). However, binary query and key representations map continuous similarities to sparse an…

cs.RO2026

ForesightSafety-VLA: A Unified Diagnostic Safety Benchmark for Vision-Language-Action Models

Mingyang Lyu, Yinqian Sun, Yiyang Jia +5

In embodied intelligence, safety is a prerequisite for reliable robot deployment in the physical world. Current vision-language-action (VLA) models continue to advance toward gener…

cs.AI2026

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Linghao Feng, Yinqian Sun, Dongqi Liang +8

Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature analysis to laboratory planning a…

cs.AI2026

Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron

Sicheng Shen, Mingyang Lv, Han Shen +7

The safety of large language models (LLMs) has increasingly emerged as a fundamental aspect of their development. Existing safety alignment for LLMs is predominantly achieved throu…

cs.NE2026

TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking Transformers

Sicheng Shen, Mingyang Lv, Bing Han +4

In recent years, Spiking Neural Networks (SNNs) have achieved remarkable progress, with Spiking Transformers emerging as a promising architecture for energy-efficient sequence mode…