4 papers
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges
Jonaid Shianifar, Blaz Mramor, Fangda Zou +5
Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distr…
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction
Jonaid Shianifar, Iias Faiud
Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different…
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
Jonaid Shianifar, Michael Schukat, Karl Mason
Multi-objective reinforcement learning (MORL) enables agents to optimize vector-valued rewards while respecting user preferences. CAPQL, a preference-conditioned actor-critic metho…
Optimizing Deep Reinforcement Learning for Adaptive Robotic Arm Control
Jonaid Shianifar, Michael Schukat, Karl Mason
In this paper, we explore the optimization of hyperparameters for the Soft Actor-Critic (SAC) and Proximal Policy Optimization (PPO) algorithms using the Tree-structured Parzen Est…