9 papers
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
Yanbo Wang, Minzheng Wang, Jian Liang +3
While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures. For safety alignment, the core ch…
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
Yanbo Wang, Yuxuan Wang, Chen Chen +6
With the wide adoption of Multimodal Models (MMs) in real-world scenarios, it is significant to efficiently train emerging MMs that exhibit increasingly complex module architecture…
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
Pengyun Zhu, Qiheng Sun, Long Wen +7
Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and complex, as well as filled with…
What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time
Dong Yan, Jian Liang, Yanbo Wang +3
Test-Time Reinforcement Learning (TTRL) enables Large Language Models (LLMs) to enhance reasoning capabilities on unlabeled test streams by deriving pseudo-rewards from majority vo…
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
Yongcan Yu, Lingxiao He, Shuo Lu +10
Recent advances in vision-language models (VLMs) reasoning have been largely attributed to the rise of reinforcement Learning (RL), which has shifted the community's focus away fro…
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
Yanbo Wang, Yongcan Yu, Jian Liang +1
The development of Long-CoT reasoning has advanced LLM performance across various tasks, including language understanding, complex problem solving, and code generation. This paradi…