Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail
He Liu, Changtao Miao, Xinjie Yang +12
Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. E…
cs.AI2026
Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
Long P. Hoang, Hai V. Le, Shaoyang Xu +2
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In…
cs.AI2026
Multilingual Fine-Tuning via Localized Gradient Conflict Resolution
Long P. Hoang, Yiran Zhao, Wei Lu +1
The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tuning these models frequently…