13 papers
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization
Hyunmin Hwang, Jaemin Kim, Choonghan Kim +2
Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to explore diverse reasoning paths. However,…
Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models
Jeongjae Lee, Jinho Chang, Jeongsol Kim +1
Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model. Although existing…
Understanding and Accelerating the Training of Masked Diffusion Language Models
Chunsan Hong, Sanghyun Lee, Chieh-Hsin Lai +5
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs are known to learn substantially more sl…
CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
Seonghyun Jin, Youngmin Kim, Sunwoo Park +1
Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing at…
Gradient-Free Noise Optimization for Reward Alignment in Generative Models
Jeongsol Kim, Hongeun Kim, Jian Wang +1
Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to deterministic generators. A natural a…
LPDP: Inference-Time Reward Control for Variable-Length DNA Generation with Edit Flows
Jeongchan Kim, Yunkyung Ko, Jong Chul Ye
We study the application of recent Edit Flows for inference-time reward control for DNA sequence generation. Unlike most reward-guided DNA generation frameworks, which operate on f…