5 papers
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
Enhao Huang, Pengyu Sun, Shuxun Wang +13
The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where software vulnerabilities and incentive misalignm…
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
Weiyi Chen, Shuaixiong Wang, Ziyun Gao +5
The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations:…
When AI reviews science: Can we trust the referee?
Jialiang Wang, Yuchen Liu, Hang Xu +7
The volume of scientific submissions continues to climb, outpacing the capacity of qualified human referees and stretching editorial timelines. At the same time, modern large langu…
DualMind: Towards Understanding Cognitive-Affective Cascades in Public Opinion Dissemination via Multi-Agent Simulation
Enhao Huang, Tongtong Pan, Shuhuai Zhang +6
Forecasting public opinion during PR crises is challenging, as existing frameworks often overlook the interaction between transient affective responses and persistent cognitive bel…
Holographic Transformers for Complex-Valued Signal Processing: Integrating Phase Interference into Self-Attention
Enhao Huang, Zhiyu Zhang, Tianxiang Xu +6
Complex-valued signals encode both amplitude and phase, yet most deep models treat attention as real-valued correlation, overlooking interference effects. We introduce the Holograp…