2 papers
cs.LG2025
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
Woosung Koh, Wonbeen Oh, Jaein Jang +7
Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (…
cs.LG2025
: Scalable Auto-Feedback for LLM-based Chart Generation
Woosung Koh, Jang Han Yoon, MinHyung Lee +7
Generating high-quality charts with Large Language Models (LLMs) presents significant challenges due to limited data and the high cost of scaling through human curation. $\langle \…