1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2026
QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models
Zhaolu Kang, Junhao Gong, Wenqing Hu +15
Large Language Models (LLMs) have shown strong capabilities across many domains, yet their evaluation in financial quantitative tasks remains fragmented and mostly limited to knowl…
cs.AI2025★ 1 cited
Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution
Zailong Tian, Zhuoheng Han, Yanzhe Chen +5
Large Language Models (LLMs) are widely used as automated judges, where practical value depends on both accuracy and trustworthy, risk-aware judgments. Existing approaches predomin…