3 papers
cs.CL2026
CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training
Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu +1
Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality fi…
cs.CL2026
Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu +1
Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filter…
cs.CV2026
PushupBench: Your VLM is not good at counting pushups
Shengzhi Li, Jiarun Chen, Karun Sharma +2
Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (…