activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2025

Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?

Divij Chawla, Ashita Bhutada, Do Duc Anh +8

We assess whether AI systems can credibly evaluate investment risk appetite-a task that must be thoroughly validated before automation. Our analysis was conducted on proprietary sy…

cs.CL2025

Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse

Maojia Song, Shang Hong Sim, Rishabh Bhardwaj +3

LLMs are an integral component of retrieval-augmented generation (RAG) systems. While many studies focus on evaluating the overall quality of end-to-end RAG systems, there is a gap…

cs.CL2025

MSTS: A Multimodal Safety Test Suite for Vision-Language Models

Paul Röttger, Giuseppe Attanasio, Felix Friedrich +19

Vision-language models (VLMs), which process image and text inputs, are increasingly integrated into chat assistants and other consumer AI applications. Without proper safeguards,…

cs.CL2024

Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability

Haonan Li, Xudong Han, Zenan Zhai +32

To address this gap, we introduce Libra-Leaderboard, a comprehensive framework designed to rank LLMs through a balanced evaluation of performance and safety. Combining a dynamic le…

cs.CL2024

Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique

Tej Deep Pala, Vernon Y. H. Toh, Rishabh Bhardwaj +1

In today's era, where large language models (LLMs) are integrated into numerous real-world applications, ensuring their safety and robustness is crucial for responsible AI usage. A…

cs.CL2024

WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models

Prannaya Gupta, Le Qi Yau, Hao Han Low +8

WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models (LLMs). It accommodates a diverse range of models, including both open-weight and…