collaborators

5 papers

cs.CL2026

Solar Open 2 Technical Report

Sungrae Park, Sanghoon Kim, Gyoungjin Gim +50

We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100B). To hold entire agent tra…

cs.LG2026

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

Hyungkyu Kang, Byeongchan Kim, Min-hwan Oh

Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-co…

cs.LG2026

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards

Deokgyu Yoon, Hyungkyu Kang, Joongkyu Lee +4

Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objective…

cs.LG2025

Adaptive Graph Learning with Transformer for Multi-Reservoir Inflow Prediction

Pengfei Hu, Ming Fan, Xiaoxue Han +5

Reservoir inflow prediction is crucial for water resource management, yet existing approaches mainly focus on single-reservoir models that ignore spatial dependencies among interco…

cs.LG2025

Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning

Hyungkyu Kang, Min-hwan Oh

In this paper, we study offline preference-based reinforcement learning (PbRL), where learning is based on pre-collected preference feedback over pairs of trajectories. While offli…