2 papers
cs.IR2025
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
Yutao Wu, Xiao Liu, Yunhao Feng +2
Large Language Models (LLMs) increasingly serve as research assistants, yet their reliability in scholarly tasks remains under-evaluated. In this work, we introduce PaperAsk, a ben…
cs.LG2025
RedTopic: Toward Topic-Diverse Red Teaming of Large Language Models
Jiale Ding, Xiang Zheng, Yutao Wu +5
As large language models (LLMs) are increasingly deployed as black-box components in real-world applications, red teaming has become essential for identifying potential risks. It t…