6 papers · 1 filter
Privacy-Preserving Text Sanitization for Distributed Agents Collaboration via Disentangled Representations
Xuan Liu, Hefeng Zhou, Sicheng Chen +6
When distributed agents exchange text across organizational boundaries, privacy leakage arises not only from explicit identifiers but also from distributional signatures such as fo…
TreeEval: Benchmark-Free Evaluation of Large Language Models through Tree Planning
Xiang Li, Yunshi Lan, Chao Yang
Recently, numerous new benchmarks have been established to evaluate the performance of large language models (LLMs) via either computing a holistic score or employing another LLM a…
Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
Zhanhui Zhou, Zhixuan Liu, Jie Liu +3
Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{we…
SEER: Facilitating Structured Reasoning and Explanation via Reinforcement Learning
Guoxin Chen, Kexin Tang, Chao Yang +3
Elucidating the reasoning process with structured explanations from question to answer is crucial, as it significantly enhances the interpretability, traceability, and trustworthin…
Inference-Time Language Model Alignment via Integrated Value Guidance
Zhixuan Liu, Zhanhui Zhou, Yuanfu Wang +2
Large language models are typically fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. In this work, we introduce $\texti…
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
Zhanhui Zhou, Jie Liu, Zhichen Dong +4
Large language models (LLMs) undergo safety alignment to ensure safe conversations with humans. However, this paper introduces a training-free attack method capable of reversing sa…