Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
Yu Feng, Chunting Zang, Chen Shen +4
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard…
cs.AI2025
Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response
Yangqing Zheng, Shunqi Mao, Dingxin Zhang +1
Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic and safety-critical environments. Howeve…