works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
collaborators

11 papers

cs.AI2026

REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment

Zhengze Huang, Luyang Yu, Di Hong +5

Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hallucinations in LRMs arise from two distinc…

cs.CR2026

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

Ruoyu Wang, Heng Zhao, Renjie Wu +4

The paper presents AgentSnare, a system that dynamically creates deceptive decoy environments to mislead and delay autonomous penetration testing agents powered by large language m…

cs.AI2026

EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning

Yitong Qiao, Lei Liu, Yue Shen +4

Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often ope…

cs.CL2026

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

Yan Wang, Zhixuan Chu, Zihao Xue +9

Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful…

cs.CL2026

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

Yuchen Yang, Wenze Lin, Enhao Huang +6

Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to spec…

cs.CV2026

JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization

Haolun Zheng, Yu He, Tailun Chen +6

Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed…