2 papers
cs.AI2026
Constrained Auto-Bidding via Generative Response Modeling
Eunseok Yang, Xingdong Zuo, Kyung-Min Kim
Auto-bidding systems aim to maximize advertiser value over long horizons under budget constraints and ratio targets such as cost-per-acquisition, yet future traffic and auction dyn…
cs.LG2025
Mitigating Suboptimality of Deterministic Policy Gradients in Complex Q-functions
Ayush Jain, Norio Kosaka, Xinhu Li +3
In reinforcement learning, off-policy actor-critic methods like DDPG and TD3 use deterministic policy gradients: the Q-function is learned from environment data, while the actor ma…