most citedGuardReasoner: Towards Reasoning-based LLM Safeguards

2 citations · 2 across the 19 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry

Shengfang Zhai, Leo Marchyok, Yuling Shi +4

Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive LMs, offering advantages such as parallel generation and bidirectional…

cs.CL2026

Beyond Hard Masks: Progressive Token Evolution for Diffusion Language Models

Linhao Zhong, Linyu Wu, Bozhen Fang +6

Diffusion Language Models (DLMs) offer a promising alternative for language modeling by enabling parallel decoding through iterative refinement. However, most DLMs rely on hard bin…

cs.CL2025

DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models

Zherui Li, Zheng Nie, Zhenhong Zhou +7

The rapid advancement of Diffusion Large Language Models (dLLMs) introduces unprecedented vulnerabilities that are fundamentally distinct from Autoregressive LLMs, stemming from th…

cs.CL2025

Efficient Reasoning via Chain of Unconscious Thought

Ruihan Gong, Yue Liu, Wenjie Qu +11

Large Reasoning Models (LRMs) achieve promising performance but compromise token efficiency due to verbose reasoning processes. Unconscious Thought Theory (UTT) posits that complex…

cs.CL2025

Efficient Inference for Large Reasoning Models: A Survey

Yue Liu, Jiaying Wu, Yufei He +11

Large Reasoning Models (LRMs) significantly improve the reasoning ability of Large Language Models (LLMs) by learning to reason, exhibiting promising performance in solving complex…

cs.CL2025

Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning

Tianyi Wu, Jingwei Ni, Bryan Hooi +5

Instruction fine-tuning (IFT) can increase the informativeness of large language models (LLMs), but may reduce their truthfulness. This trade-off arises because IFT steers LLMs to…