2 citations · 3 across the 20 of their papers we have counts for
22 papers
One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread
Zhuoran Lu, Weilong Wang, Yangyang Yu +4
Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-making, these indicators are embedde…
NTDH: Complex Reasoning for Comprehensive Affective Analysis
Tianlei Zhu, Zhiwei Liu, Yuyan Wang +2
Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is…
Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
Hanhua Hong, Yizhi Li, Jiaoyan Chen +4
Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repositor…
Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs
Polydoros Giannouris, Mohsinul Kabir, Sophia Ananiadou
LLM deception is often evaluated through direct markers such as fabricated claims, explicit lies, or strategic concealment. However, many real-world misleading communications do no…
AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements
Zhiwei Liu, Yueru He, Qing Ou +4
Large language models (LLMs) have shown strong performance in financial analysis and surface-level factual error detection, yet their ability to identify fraudulent financial misin…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…