3 papers
cs.CL2025
Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
Chenghao Zhu, Meiling Tao, Tiannan Wang +3
Faithfully personalizing large language models (LLMs) to align with individual user preferences is a critical but challenging task. While supervised fine-tuning (SFT) quickly reach…
cs.CL2025
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
Dongyi Ding, Tiannan Wang, Chenghao Zhu +3
Large language models (LLMs) excel at reasoning tasks requiring long thought sequences for planning, reflection, and refinement. However, their substantial model size and high comp…
cs.CL2025
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
Meiling Tao, Chenghao Zhu, Dongyi Ding +3
With the rapid improvement in the general capabilities of LLMs, LLM personalization, i.e., how to build LLM systems that can generate personalized responses or services that are ta…