Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Reinforced Strategy Optimization for Conversational Recommender Systems via Network-of-Experts
Xiaoyan Zhao, Ming Yan, Yang Zhang +6
Conversational Recommender Systems (CRSs) aim to provide personalized recommendations through multi-turn natural language interactions with users. Given the strong interaction and…
cs.CL2025
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
Zhiling Ye, Yun Yue, Haowen Wang +11
Open-ended evaluation is essential for deploying large language models in real-world settings. In studying HealthBench, we observe that using the model itself as a grader and gener…
cs.CL2024
RuleAlign: Making Large Language Models Better Physicians with Diagnostic Rule Alignment
Xiaohan Wang, Xiaoyan Yang, Yuqi Zhu +7
Large Language Models (LLMs) like GPT-4, MedPaLM-2, and Med-Gemini achieve performance competitively with human experts across various medical benchmarks. However, they still face…