1 citations · 1 across the 5 of their papers we have counts for
9 papers
How Reliable are Confidence Estimators for Large Reasoning Models? A Systematic Benchmark on High-Stakes Domains
Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur +4
The miscalibration of Large Reasoning Models (LRMs) undermines their reliability in high-stakes domains, necessitating methods to accurately estimate the confidence of their long-f…
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
Elizabeth Fons, Elena Kochkina, Rachneet Kaur +5
This paper explores the potential of large language models (LLMs) to generate financial reports from time series data. We propose a framework encompassing prompt engineering, model…
A Variational Approach for Mitigating Entity Bias in Relation Extraction
Samuel Mensah, Elena Kochkina, Jabez Magomere +3
Mitigating entity bias is a critical challenge in Relation Extraction (RE), where models often rely excessively on entities, resulting in poor generalization. This paper presents a…
GenPlanX. Generation of Plans and Execution
Daniel Borrajo, Giuseppe Canonaco, Tomás de la Rosa +10
Classical AI Planning techniques generate sequences of actions for complex tasks. However, they lack the ability to understand planning tasks when provided using natural language.…
Conservative Bias in Large Language Models: Measuring Relation Predictions
Toyin Aguda, Erik Wilson, Allan Anzagira +2
Large language models (LLMs) exhibit pronounced conservative bias in relation extraction tasks, frequently defaulting to No_Relation label when an appropriate option is unavailable…
FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and Challenging
Zichen Tang, Haihong E, Ziyan Ma +10
We introduce FinanceReasoning, a novel benchmark designed to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. Compare…