4 papers
K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean
Minkyeong Jeon, Hyemin Jeong, Yerang Kim +3
Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification mod…
Semi-gradient DICE for Offline Constrained Reinforcement Learning
Woosung Kim, JunHo Seo, Jongmin Lee +1
Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliabl…
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
Hosung Lee, Sejin Kim, Seungpil Lee +4
This paper introduces ARCLE, an environment designed to facilitate reinforcement learning research on the Abstraction and Reasoning Corpus (ARC). Addressing this inductive reasonin…
Offline Imitation Learning by Controlling the Effective Planning Horizon
Hee-Jun Ahn, Seong-Woong Shim, Byung-Jun Lee
In offline imitation learning (IL), we generally assume only a handful of expert trajectories and a supplementary offline dataset from suboptimal behaviors to learn the expert poli…