7 papers
Soft Token Alignment for Cross-Lingual Reasoning
Jiayi He, Jungsoo Park, Wei Xu +1
Multilingual large language models often produce inconsistent reasoning and answers for semantically equivalent prompts in different languages. Prior work suggests that intermediat…
Making Expert Reasoning Learnable with Self-Distillation
Ethan Mendes, Jungsoo Park, Alan Ritter
Improving the reasoning capabilities of large language models (LLMs) typically relies either on the model's ability to sample a correct solution to be reinforced or the existence o…
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
Jungsoo Park, Hyungjoo Chae, Ethan Mendes +4
Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training objectives score each decoded floati…
Safe and Scalable Web Agent Learning via Recreated Websites
Hyungjoo Chae, Jungsoo Park, Alan Ritter
Training autonomous web agents is fundamentally limited by the environments they learn from: real-world websites are unsafe to explore, hard to reset, and rarely provide verifiable…
Anticipatory Evaluation of Language Models
Jungsoo Park, Ethan Mendes, Gabriel Stanovsky +1
Progress in large language models is increasingly constrained by an evaluation bottleneck: benchmarks must be built and models run before iteration can begin. We investigate whethe…
Data Transformation Strategies to Remove Heterogeneity
Sangbong Yoo, Jaeyoung Lee, Chanyoung Yoon +8
Data heterogeneity is a prevalent issue, stemming from various conflicting factors, making its utilization complex. This uncertainty, particularly resulting from disparities in dat…