7 papers
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
Enze Ma, Yufan Zhou, Wei-Chieh Huang +7
Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms. In practice, howeve…
Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits
Huanhuan Ma, Haisong Gong, Xiaoyuan Yi +3
As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding their behavioral characteristics becomes esse…
Training Data for Large Language Model
Yiming Ju, Huanhuan Ma
In 2022, with the release of ChatGPT, large-scale language models gained widespread attention. ChatGPT not only surpassed previous models in terms of parameters and the scale of it…
Does Knowledge Localization Hold True? Surprising Differences Between Entity and Relation Perspectives in Language Models
Yifan Wei, Xiaoyan Yu, Yixuan Weng +4
Large language models encapsulate knowledge and have demonstrated superior performance on various natural language processing tasks. Recent studies have localized this knowledge to…
Navigating the Noisy Crowd: Finding Key Information for Claim Verification
Haisong Gong, Huanhuan Ma, Qiang Liu +2
Claim verification is a task that involves assessing the truthfulness of a given claim based on multiple evidence pieces. Using large language models (LLMs) for claim verification…
EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification
Huanhuan Ma, Weizhi Xu, Yifan Wei +4
Fact verification aims to automatically probe the veracity of a claim based on several pieces of evidence. Existing works are always engaging in accuracy improvement, let alone exp…