most citedFinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking

1 citations · 1 across the 5 of their papers we have counts for

collaborators

9 papers

cs.CL2026

How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains

Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur +4

The miscalibration of Large Reasoning Models (LRMs) undermines their reliability in high-stakes domains, necessitating methods to accurately estimate the confidence of their long-f…

cs.CL2025

AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation

Elizabeth Fons, Elena Kochkina, Rachneet Kaur +5

This paper explores the potential of large language models (LLMs) to generate financial reports from time series data. We propose a framework encompassing prompt engineering, model…

cs.CL2025

A Variational Approach for Mitigating Entity Bias in Relation Extraction

Samuel Mensah, Elena Kochkina, Jabez Magomere +3

Mitigating entity bias is a critical challenge in Relation Extraction (RE), where models often rely excessively on entities, resulting in poor generalization. This paper presents a…

cs.AI2025

GenPlanX. Generation of Plans and Execution

Daniel Borrajo, Giuseppe Canonaco, Tomás de la Rosa +10

Classical AI Planning techniques generate sequences of actions for complex tasks. However, they lack the ability to understand planning tasks when provided using natural language.…

cs.CL2025

Conservative Bias in Large Language Models: Measuring Relation Predictions

Toyin Aguda, Erik Wilson, Allan Anzagira +2

Large language models (LLMs) exhibit pronounced conservative bias in relation extraction tasks, frequently defaulting to No_Relation label when an appropriate option is unavailable…

cs.CL2025

FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging

Zichen Tang, Haihong E, Ziyan Ma +10

We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compare…