3 papers
cs.LG2026
Rollout-Level Advantage-Prioritized Experience Replay for GRPO
Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang +2
Reinforcement learning from verifiable rewards with GRPO is a standard approach for post-training reasoning LLMs. It remains sample inefficient. Each rollout is used for a single g…
cs.SE2025
Automating Code Generation for Semiconductor Equipment Control from Developer Utterances with LLMs
Youngkyoung Kim, Sanghyeok Park, Misoo Kim +3
Semiconductors form the backbone of modern electronics, with their manufacturing and testing relying on highly specialized equipment and domain-specific programming languages. Equi…
cs.CV2025
ODPG: Outfitting Diffusion with Pose Guided Condition
Seohyun Lee, Jintae Park, Sanghyeok Park
Virtual Try-On (VTON) technology allows users to visualize how clothes would look on them without physically trying them on, gaining traction with the rise of digitalization and on…