2 papers
cs.CY2025
Ask What Your Country Can Do For You: Towards a Public Red Teaming Model
Wm. Matthew Kennedy, Cigdem Patlak, Jayraj Dave +10
AI systems have the potential to produce both benefits and harms, but without rigorous and ongoing adversarial evaluation, AI actors will struggle to assess the breadth and magnitu…
cs.AI2025
A Methodology for Assessing the Risk of Metric Failure in LLMs Within the Financial Domain
William Flanagan, Mukunda Das, Rajitha Ramanayake +10
As Generative Artificial Intelligence is adopted across the financial services industry, a significant barrier to adoption and usage is measuring model performance. Historical mach…