3 papers
cs.RO2026
RobotValues: Evaluating Household Robots When Human Values Conflict
Jongwook Han, Hyeongjin Kim, Yohan Jo
While household robots are often evaluated based on task completion, everyday domestic environments involve value-conflicting situations in which robots are expected to choose acti…
cs.LG2025
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
Woosung Koh, Wonbeen Oh, Jaein Jang +7
Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (…
cs.LG2025
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
Woosung Koh, Wonbeen Oh, Siyeol Kim +5
Multi-agent reinforcement learning has demonstrated significant potential in addressing complex cooperative tasks across various real-world applications. However, existing MARL app…