Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training
Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu +1
Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality fi…
cs.CL2026
Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu +1
Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filter…