7 papers
PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing
Keshen Zhou, Lintao Wang, Suqin Yuan +3
The paper introduces PaperRouter-Agent, a training-free LLM-based system that routes new research papers into a user’s personal folder hierarchy by examining the contents of existi…
Understanding Diversity Collapse in RLVR via the Lens of Overtraining
Suqin Yuan, Jinkun Chen, Jiyang Zheng +6
Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language models. However, RLVR often suffers from \em…
FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
Runqi Lin, Alasdair Paren, Suqin Yuan +4
The integration of new modalities enhances the capabilities of multimodal large language models (MLLMs) but also introduces additional vulnerabilities. In particular, simple visual…
Mitigating Mismatch within Reference-based Preference Optimization
Suqin Yuan, Xingrui Yu, Jiyang Zheng +4
Direct Preference Optimization (DPO) has become the de facto standard for offline preference alignment of large language models, but its reliance on a reference policy introduces a…
Enhancing Sample Selection Against Label Noise by Cutting Mislabeled Easy Examples
Suqin Yuan, Lei Feng, Bo Han +1
Sample selection is a prevalent approach in learning with noisy labels, aiming to identify confident samples for training. Although existing sample selection methods have achieved…
Early Stopping Against Label Noise Without Validation Data
Suqin Yuan, Lei Feng, Tongliang Liu
Early stopping methods in deep learning face the challenge of balancing the volume of training and validation data, especially in the presence of label noise. Concretely, sparing m…