2 papers
q-fin.PM2025
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
Daniil Karzanov, Rubén Garzón, Mikhail Terekhov +3
This paper introduces a novel agent-based approach for enhancing existing portfolio strategies using Proximal Policy Optimization (PPO). Rather than focusing solely on traditional…
cs.LG2025
Contextual Bandit Optimization with Pre-Trained Neural Networks
Mikhail Terekhov
Bandit optimization is a difficult problem, especially if the reward model is high-dimensional. When rewards are modeled by neural networks, sublinear regret has only been shown un…