3 papers
cs.AI2026
ATLAS: A Multi-LLM Training Framework for EvoDPO with Adaptive Reference Evolution
Ujin Jeon, Jiyong Kwon, Madison Ann Sullivan +2
Recent multi-LLM agent systems have shown promising capabilities for automated problem-solving, yet they predominantly rely on frozen agents or static fine-tuning pipelines. To add…
cs.LG2025
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
Weixin Wang, Haoyang Zheng, Guang Lin +2
Most existing approximate Thompson Sampling (TS) algorithms for multi-armed bandits use Stochastic Gradient Langevin Dynamics (SGLD) or its variants in each round to sample from th…
cs.LG2024
Constrained Exploration via Reflected Replica Exchange Stochastic Gradient Langevin Dynamics
Haoyang Zheng, Hengrong Du, Qi Feng +2
Replica exchange stochastic gradient Langevin dynamics (reSGLD) is an effective sampler for non-convex learning in large-scale datasets. However, the simulation may encounter stagn…