2 papers
cs.CL2026
Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation
Sarra Gharsallah, Adele Robaldo, Mariia Tokareva +5
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuali…
cs.LG2022
Optimizing Revenue Maximization and Demand Learning in Airline Revenue Management
Giovanni Gatti Pinheiro, Michael Defoin-Platel, Jean-Charles Regin
Correctly estimating how demand respond to prices is fundamental for airlines willing to optimize their pricing policy. Under some conditions, these policies, while aiming at maxim…