4 papers
Benchmarking Real-Time Question Answering via Executable Code Workflows
Wenjie Zhou, Yuan Gao, Xin Zhou +5
Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and ther…
Adversarial Alignment: Ensuring Value Consistency in Large Language Models for Sensitive Domains
Yuan Gao, Zhigang Liu, Xinyu Yao +2
With the wide application of large language models (LLMs), the problems of bias and value inconsistency in sensitive domains have gradually emerged, especially in terms of race, so…
MSME: A Multi-Stage Multi-Expert Framework for Zero-Shot Stance Detection
Yuanshuo Zhang, Aohua Li, Bo Chen +2
LLM-based approaches have recently achieved impressive results in zero-shot stance detection. However, they still struggle in complex real-world scenarios, where stance understandi…
WenyanGPT: A Large Language Model for Classical Chinese Tasks
Xinyu Yao, Mengdi Wang, Bo Chen +1
Classical Chinese, as the core carrier of Chinese culture, plays a crucial role in the inheritance and study of ancient literature. However, existing natural language processing mo…