1 citations · 1 across the 8 of their papers we have counts for
10 papers
Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
Xiaotian Zhang, Yuan Wang, Ruizhe Chen +3
The deployment of Large Language Models (LLMs) in interactive systems necessitates a deep alignment with the nuanced and dynamic preferences of individual users. Current alignment…
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
Ruizhe Chen, Zhiting Fan, Tianze Luo +7
Video Temporal Grounding (VTG) aims to localize relevant temporal segments in videos given natural language queries. Despite recent progress with large vision-language models (LVLM…
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning
Xiaotian Zhang, Yuan Wang, Zhaopeng Feng +6
Medical Question-Answering (QA) encompasses a broad spectrum of tasks, including multiple choice questions (MCQ), open-ended text generation, and complex computational reasoning. D…
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making
Songtao Jiang, Yuan Wang, Ruizhe Chen +8
In medical visual question answering (Med-VQA), achieving accurate responses relies on three critical steps: precise perception of medical imaging data, logical reasoning grounded…
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
Zhaopeng Feng, Shaosheng Cao, Jiahan Ren +7
Large-scale reinforcement learning (RL) methods have proven highly effective in enhancing the reasoning abilities of large language models (LLMs), particularly for tasks with verif…
BiasGuard: A Reasoning-enhanced Bias Detection Tool For Large Language Models
Zhiting Fan, Ruizhe Chen, Zuozhu Liu
Identifying bias in LLM-generated content is a crucial prerequisite for ensuring fairness in LLMs. Existing methods, such as fairness classifiers and LLM-based judges, face limitat…