2 papers
cs.CL2025
Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?
Divij Chawla, Ashita Bhutada, Do Duc Anh +8
We assess whether AI systems can credibly evaluate investment risk appetite-a task that must be thoroughly validated before automation. Our analysis was conducted on proprietary sy…
cs.CL2024
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
Prannaya Gupta, Le Qi Yau, Hao Han Low +8
WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models (LLMs). It accommodates a diverse range of models, including both open-weight and…