4 papers
OViP: Online Vision-Language Preference Learning for VLM Hallucination
Shujun Liu, Siyuan Wang, Zejun Li +3
Large vision-language models (LVLMs) remain vulnerable to hallucination, often generating content misaligned with visual inputs. Although recent training-based approaches aim to mi…
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
Zhuohang Long, Siyuan Wang, Shujun Liu +3
Jailbreak attacks, where harmful prompts bypass generative models' built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed,…
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
Shengbin Yue, Ting Huang, Zheng Jia +5
Large Language Models (LLMs) have significantly advanced legal intelligence, but the scarcity of scenario data impedes the progress toward interactive legal scenarios. This paper i…
HAF-RM: A Hybrid Alignment Framework for Reward Model Training
Shujun Liu, Xiaoyu Shen, Yuhang Lai +5
The reward model has become increasingly important in alignment, assessment, and data construction for large language models (LLMs). Most existing researchers focus on enhancing re…