4 papers
CrossAlpha: An Annual-Report Benchmark for Cross-Market Factor Research (with LLM Agents)
Qian Wang, Zhongyi Tong, Nuo Chen +2
Cross-market factor research studies whether firm-level signals from one or more markets can predict returns in a target market, but existing public benchmarks do not support cross…
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
Chengzhi Liu, Yuzhe Yang, Sophia Xiao Pu +14
Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stal…
Judging with Personality and Confidence: A Study on Personality-Conditioned LLM Relevance Assessment
Nuo Chen, Hanpei Fang, Piaohong Wang +3
Recent studies have shown that prompting can enable large language models (LLMs) to simulate specific personality traits and produce behaviors that align with those traits. However…
Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing
Nuo Chen, Hanpei Fang, Jiqun Liu +3
Recent research has explored LLMs as scalable tools for relevance labeling, but studies indicate they are susceptible to priming effects, where prior relevance judgments influence…