From the 1 of 5 linked papers with an AI index.
5 papers
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang +23
The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation
Andreas Haupt, Justin Hartenstein, Anka Reuel +2
AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has received far less attention:…
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
Max Lamparth, Daniel Fein, Andreas Haupt +2
Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than…
General Preference Reinforcement Learning
Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal +5
Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Online reinforcement learning (RL) with verifiable rewards drives emergent reason…
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning
Ian Gemp, Andreas Haupt, Luke Marris +2
Behavioral diversity, expert imitation, fairness, safety goals and others give rise to preferences in sequential decision making domains that do not decompose additively across tim…