From the 1 of 11 linked papers with an AI index.
11 papers
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
Zhengze Huang, Luyang Yu, Di Hong +5
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…
AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
Ruoyu Wang, Heng Zhao, Renjie Wu +4
The paper presents AgentSnare, a system that dynamically creates deceptive decoy environments to mislead and delay autonomous penetration testing agents powered by large language m…
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
Yitong Qiao, Lei Liu, Yue Shen +4
Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often ope…
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
Yan Wang, Zhixuan Chu, Zihao Xue +9
Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful…
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
Yuchen Yang, Wenze Lin, Enhao Huang +6
Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to spec…
JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization
Haolun Zheng, Yu He, Tailun Chen +6
Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed…