7 papers
FIDES: Faithful Inference via Deep Evidence Signals for Retrieval-Memory Conflict in RAG
Zhe Yu, Wenpeng Xing, Tiancheng Zhao +3
When retrieved evidence contradicts parametric memory, language models frequently ignore context and default to memorized priors -- a failure that undermines the core purpose of re…
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
Jin Li, Zhebo Wang, Tianliang Lu +3
Entropy-based inference methods have gained traction for improving the reliability of Large Language Models (LLMs). However, many existing approaches, such as entropy minimization…
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
Zhebo Wang, Xiaohu Mu, Zijie Zhou +4
Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, p…
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
Wenpeng Xing, Mohan Li, Bohan Yang +5
Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. We introduce Latent Fusion Jailbreak (LFJ)…
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
Wenpeng Xing, Lanyi Wei, Haixiao Hu +5
The rapid proliferation of large language models (LLMs) in applications targeting children and adolescents necessitates a fundamental reassessment of prevailing AI safety framework…
PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement
Xubin Yue, Zhenhua Xu, Wenpeng Xing +3
Addressing the intellectual property protection challenges in commercial deployment of large language models (LLMs), existing black-box fingerprinting techniques face dual challeng…