controlled markov process 1information theory 1mutual information 1optimal policy 1reinforcement learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
Calculating Mutual Information between a Reward Maximizer and its Environment
Alfred Harwood, Jose Faustino, Alex Altair
The paper derives an exact information-theoretic bound showing that an optimal deterministic policy in a Controlled Markov Process reveals n log m bits about the environment, quant…
cs.AI2025
Approximating Human Preferences Using a Multi-Judge Learned System
Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…