15 papers
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali
As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14…
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
Fabrizio Dimino, Abhinav Arun, Bhaskarjit Sarmah +1
Large language models (LLMs) are increasingly being used to extract structured knowledge from unstructured financial text. Although prior studies have explored various extraction m…
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
Mahesh Kumar, Bhaskarjit Sarmah, Stefano Pasquali
As organizations increasingly integrate AI-powered question-answering systems into financial information systems for compliance, risk assessment, and decision support, ensuring the…
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali
The rapid adoption of large language models (LLMs) in financial services introduces new operational, regulatory, and security risks. Yet most red-teaming benchmarks remain domain-a…
Deep Reinforcement Learning for Optimum Order Execution: Mitigating Risk and Maximizing Returns
Khabbab Zakaria, Jayapaulraj Jerinsh, Andreas Maier +3
Optimal Order Execution is a well-established problem in finance that pertains to the flawless execution of a trade (buy or sell) for a given volume within a specified time frame.…
Uncovering Representation Bias for Investment Decisions in Open-Source Large Language Models
Fabrizio Dimino, Krati Saxena, Bhaskarjit Sarmah +1
Large Language Models are increasingly adopted in financial applications to support investment workflows. However, prior studies have seldom examined how these models reflect biase…