activity
20242026
collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Fengxian Ji, Yuke Li, Jingpu Yang +8

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this que…

cs.CL2026

ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

Jie Gong, Maowei Jiang, Zhiwei Liu +14

Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily asses…

cs.CL2026

MisSpans: Fine-Grained False Span Identification in Cross-Domain Fake News

Zhiwei Liu, Paul Thompson, Jiaqi Rong +5

Online misinformation is increasingly pervasive, yet most existing benchmarks and methods evaluate veracity at the level of whole claims or paragraphs using coarse binary labels, o…

cs.CL2026

RAAR: Retrieval Augmented Agentic Reasoning for Cross-Domain Misinformation Detection

Zhiwei Liu, Runteng Guo, Baojie Qu +4

Cross-domain misinformation detection is challenging, as misinformation arises across domains with substantial differences in knowledge and discourse. Existing methods often rely o…

cs.CL2025

MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy

Mengxi Xiao, Kailai Yang, Pengde Zhao +12

Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into…

cs.CL2025

DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation

Enze Zhang, Jiaying Wang, Mengxi Xiao +7

Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on su…