6 papers
ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
Jie Gong, Maowei Jiang, Zhiwei Liu +14
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily asses…
Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs
Jun-Yu Pan, Yansen Wang, Enze Zhang +3
Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce,…
HarmThoughts: A Benchmark for Fine-Grained Harmful Behavior Detection in Reasoning Traces
Ishita Kakkar, Enze Zhang, Rheeya Uppaal +1
Large reasoning models (LRMs) produce complex, multi-step reasoning traces, yet safety evaluation remains focused on final outputs, overlooking how harm emerges during reasoning. W…
Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
Yuechen Jiang, Enze Zhang, Md Mohsinul Kabir +4
Recent advances in vision-language models (VLMs) have improved image captioning for cultural heritage. However, inferring structured cultural metadata (e.g., creator, origin, perio…
MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy
Mengxi Xiao, Kailai Yang, Pengde Zhao +12
Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, and often ambiguous evidence into…
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
Enze Zhang, Jiaying Wang, Mengxi Xiao +7
Large language models (LLMs) have substantially advanced machine translation (MT), yet their effectiveness in translating web novels remains unclear. Existing benchmarks rely on su…