activity
20242026
collaborators

13 papers

cs.CL2026

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

Reetu Raj Harsh, Bhaskarjit Sarmah, Stefano Pasquali

As Large Language Models (LLMs) are increasingly deployed in safety-critical applications, robust content moderation becomes essential. We present a comprehensive evaluation of 14…

q-fin.CP2026

FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation

Fabrizio Dimino, Abhinav Arun, Bhaskarjit Sarmah +1

Large language models (LLMs) are increasingly being used to extract structured knowledge from unstructured financial text. Although prior studies have explored various extraction m…

cs.CL2026

FinReflectKG -- HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems

Mahesh Kumar, Bhaskarjit Sarmah, Stefano Pasquali

As organizations increasingly integrate AI-powered question-answering systems into financial information systems for compliance, risk assessment, and decision support, ensuring the…

q-fin.CP2026

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali

The rapid adoption of large language models (LLMs) in financial services introduces new operational, regulatory, and security risks. Yet most red-teaming benchmarks remain domain-a…

cs.AI2025

Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI

Samarth Sarin, Lovepreet Singh, Bhaskarjit Sarmah +1

Agentic memory is emerging as a key enabler for large language models (LLM) to maintain continuity, personalization, and long-term context in extended user interactions, critical c…

q-fin.CP2025

Uncovering Representation Bias for Investment Decisions in Open-Source Large Language Models

Fabrizio Dimino, Krati Saxena, Bhaskarjit Sarmah +1

Large Language Models are increasingly adopted in financial applications to support investment workflows. However, prior studies have seldom examined how these models reflect biase…