10 citations · 11 across the 7 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
Aladin Djuhera, Farhan Ahmed, Swanand Ravindra Kadhe +3
Aligning large language models (LLMs) is a central objective of post-training, often achieved through reward modeling and reinforcement learning methods. Among these, direct prefer…
cs.CL2025
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
Aladin Djuhera, Swanand Ravindra Kadhe, Syed Zawad +3
Recent work on large language models (LLMs) has increasingly focused on post-training and alignment with datasets curated to enhance instruction following, world knowledge, and spe…