13 papers
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali
As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14…
FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
Fabrizio Dimino, Abhinav Arun, Bhaskarjit Sarmah +1
Large language models (LLMs) are increasingly being used to extract structured knowledge from unstructured financial text. Although prior studies have explored various extraction m…
FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems
Mahesh Kumar, Bhaskarjit Sarmah, Stefano Pasquali
As organizations increasingly integrate AI-powered question-answering systems into financial information systems for compliance, risk assessment, and decision support, ensuring the…
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali
The rapid adoption of large language models (LLMs) in financial services introduces new operational, regulatory, and security risks. Yet most red-teaming benchmarks remain domain-a…
Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI
Samarth Sarin, Lovepreet Singh, Bhaskarjit Sarmah +1
Agentic memory is emerging as a key enabler for large language models (LLM) to maintain continuity, personalization, and long-term context in extended user interactions, critical c…
Uncovering Representation Bias for Investment Decisions in Open-Source Large Language Models
Fabrizio Dimino, Krati Saxena, Bhaskarjit Sarmah +1
Large Language Models are increasingly adopted in financial applications to support investment workflows. However, prior studies have seldom examined how these models reflect biase…