activity
20182026
most citedPAC-Bayesian Lifelong Learning For Multi-Armed Bandits

8 citations · 18 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

23 papers · 1 filter

cs.LG2026

Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies

Magnus Victor Boock, Abdullah Akgül, Mustafa Mert Çelikok +1

We present a new operator-theoretic representation learning framework for offline reinforcement learning that recovers the directed temporal geometry of a controlled Markov process…

cs.LG2026

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration

Manuel Haussmann, Mustafa Mert Çelikok, Melih Kandemir

While reinforcement learning (RL) promises to revolutionize the control of complex nonlinear robotic systems, a profound gap persists between the heuristic success of model-free of…

cs.LG2026

Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards

Orhun Bugra Baran, Melih Kandemir, Ramazan Gokberk Cinbis

Autoregressive (AR) models are highly effective for image generation, yet their standard maximum-likelihood estimation training lacks direct optimization for sample quality and div…

cs.LG2026

Distributional Active Inference

Abdullah Akgül, Gulcin Baykal, Manuel Haußmann +2

Optimal control of complex environments with robotic systems faces two complementary and intertwined challenges: efficient organization of sensory state information and far-sighted…

cs.LG2025

ObjectRL: An Object-Oriented Reinforcement Learning Codebase

Gulcin Baykal, Abdullah Akgül, Manuel Haussmann +4

ObjectRL is an open-source Python codebase for deep reinforcement learning (RL), designed for research-oriented prototyping with minimal programming effort. Unlike existing codebas…

cs.LG2025

Adaptive Ensemble Aggregation for Actor-Critics

Nicklas Werge, Yi-Shan Wu, Manuel Haussmann +2

Ensembles are ubiquitous in off-policy actor-critic learning, yet their efficacy depends critically on how they are aggregated. Current methods typically rely on static rules or ta…