activity
20242026
collaborators

15 papers

cs.CL2026

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali

As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14…

q-fin.CP2026

FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation

Fabrizio Dimino, Abhinav Arun, Bhaskarjit Sarmah +1

Large language models (LLMs) are increasingly being used to extract structured knowledge from unstructured financial text. Although prior studies have explored various extraction m…

cs.CL2026

FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems

Mahesh Kumar, Bhaskarjit Sarmah, Stefano Pasquali

As organizations increasingly integrate AI-powered question-answering systems into financial information systems for compliance, risk assessment, and decision support, ensuring the…

q-fin.CP2026

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali

The rapid adoption of large language models (LLMs) in financial services introduces new operational, regulatory, and security risks. Yet most red-teaming benchmarks remain domain-a…

q-fin.CP2026

Deep Reinforcement Learning for Optimum Order Execution: Mitigating Risk and Maximizing Returns

Khabbab Zakaria, Jayapaulraj Jerinsh, Andreas Maier +3

Optimal Order Execution is a well-established problem in finance that pertains to the flawless execution of a trade (buy or sell) for a given volume within a specified time frame.…

q-fin.CP2025

Uncovering Representation Bias for Investment Decisions in Open-Source Large Language Models

Fabrizio Dimino, Krati Saxena, Bhaskarjit Sarmah +1

Large Language Models are increasingly adopted in financial applications to support investment workflows. However, prior studies have seldom examined how these models reflect biase…