From the 1 of 21 linked papers with an AI index.
21 papers
NTDH: Complex Reasoning for Comprehensive Affective Analysis
Tianlei Zhu, Zhiwei Liu, Yuyan Wang +2
Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is…
Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
Hanhua Hong, Yizhi Li, Jiaoyan Chen +4
The paper conducts a meta‑evaluation of rubrics generated by large language models for assessing the reproducibility of research papers, comparing intrinsic semantic similarity and…
Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs
Polydoros Giannouris, Mohsinul Kabir, Sophia Ananiadou
LLM deception is often evaluated through direct markers such as fabricated claims, explicit lies, or strategic concealment. However, many real-world misleading communications do no…
AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements
Zhiwei Liu, Yueru He, Qing Ou +4
Large language models (LLMs) have shown strong performance in financial analysis and surface-level factual error detection, yet their ability to identify fraudulent financial misin…
XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad
Mohsinul Kabir, Tasnim Ahmed, Md Mezbaur Rahman +5
Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cultural contexts. However, progress in e…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…