Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
Leitian Tao, Xuefeng Du, Sharon Li
Reward modeling, crucial for aligning large language models (LLMs) with human preferences, is often bottlenecked by the high cost of preference data. Existing textual data synthesi…
cs.CL2025
Challenges and Future Directions of Data-Centric AI Alignment
Min-Hsuan Yeh, Jeffrey Wang, Xuefeng Du +4
As AI systems become increasingly capable and influential, ensuring their alignment with human values, preferences, and goals has become a critical research focus. Current alignmen…
cs.CL2024
Safety-Aware Fine-Tuning of Large Language Models
Hyeong Kyu Choi, Xuefeng Du, Yixuan Li
Fine-tuning Large Language Models (LLMs) has emerged as a common practice for tailoring models to individual needs and preferences. The choice of datasets for fine-tuning can be di…