2 papers
cs.AI2026
Seven simple steps for log analysis in AI systems
Magda Dubois, Ekin Zorer, Maia Hamin +17
AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess…
cs.LG2026
Skewed Score: A statistical framework to assess autograders
Magda Dubois, Harry Coppock, Mario Giulianelli +3
The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a…