14 papers
Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
Orhun Bugra Baran, Melih Kandemir, Ramazan Gokberk Cinbis
Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimization for sample quality and div…
Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies
Magnus Victor Boock, Abdullah Akgül, Mustafa Mert Ãelikok +1
We present a new operator-theoretic representation learning framework for offline reinforcement learning that recovers the directed temporal geometry of a controlled Markov process…
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration
Manuel Haussmann, Mustafa Mert Ãelikok, Melih Kandemir
While reinforcement learning (RL) promises to revolutionize the control of complex nonlinear robotic systems, a profound gap persists between the heuristic success of model-free of…
Adaptive Ensemble Aggregation for Actor-Critics
Nicklas Werge, Yi-Shan Wu, Manuel Haussmann +2
Ensembles are ubiquitous in off-policy actor-critic learning, yet their efficacy depends critically on how they are aggregated. Current methods typically rely on static rules or ta…
Distributional Active Inference
Abdullah Akgül, Abdullah Akgül, Gulcin Baykal +5
Optimal control of complex environments with robotic systems faces two complementary and intertwined challenges: efficient organization of sensory state information and far-sighted…
Deep Actor-Critics with Tight Risk Certificates
Bahareh Tasdighi, Manuel Haussmann, Yi-Shan Wu +2
Deep actor-critic algorithms have reached a level where they influence everyday life. They are a driving force behind continual improvement of large language models through user fe…