activity
20122026
most citedExploiting Submodular Value Functions For Scaling Up Active Perception

22 citations · 83 across the 32 of their papers we have counts for

collaborators

42 papers

cs.MA2026

Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response

Ariyan Bighashdel, Thiago D. Simão, Frans A. Oliehoek

Multi-agent reinforcement learning (MARL) offers a scalable alternative to exact game-theoretic analysis but suffers from non-stationarity and the need to maintain diverse populati…

cs.LG2025

Learning to Focus: Prioritizing Informative Histories with Structured Attention Mechanisms in Partially Observable Reinforcement Learning

Daniel De Dios Allegue, Jinke He, Frans A. Oliehoek

Transformers have shown strong ability to model long-term dependencies and are increasingly adopted as world models in model-based reinforcement learning (RL) under partial observa…

cs.LG2025

Conditional Policy Generator for Dynamic Constraint Satisfaction and Optimization

Wook Lee, Frans A. Oliehoek

Leveraging machine learning methods to solve constraint satisfaction problems has shown promising, but they are mostly limited to a static situation where the problem description i…

cs.LG2025

Timing the Match: A Deep Reinforcement Learning Approach for Ride-Hailing and Ride-Pooling Services

Yiman Bao, Jie Gao, Jinke He +2

Efficient timing in ride-matching is crucial for improving the performance of ride-hailing and ride-pooling services, as it determines the number of drivers and passengers consider…

cs.AI2025

SHARPIE: A Modular Framework for Reinforcement Learning and Human-AI Interaction Experiments

Hüseyin Aydın, Kevin Godin-Dubois, Libio Goncalvez Braz +6

Reinforcement learning (RL) offers a general approach for modeling and training AI agents, including human-AI interaction scenarios. In this paper, we propose SHARPIE (Shared Human…

cs.LG2024

SimuDICE: Offline Policy Optimization Through World Model Updates and DICE Estimation

Catalin E. Brita, Stephan Bongers, Frans A. Oliehoek

In offline reinforcement learning, deriving an effective policy from a pre-collected set of experiences is challenging due to the distribution mismatch between the target policy an…