5 papers
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
Junpeng Ding, Zichen Tang, Haihong E +17
We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from…
AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images
Bo Zhang, Tzu-Yen Ma, Zichen Tang +18
We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) D…
Opportunities and Challenges of Large Language Models for Low-Resource Languages in Humanities Research
Tianyang Zhong, Zhenyuan Yang, Zhengliang Liu +11
Low-resource languages serve as invaluable repositories of human history, embodying cultural evolution and intellectual diversity. Despite their significance, these languages face…
FNF: Functional Network Fingerprint for Large Language Models
Yiheng Liu, Junhao Ning, Sichen Xia +8
The development of large language models (LLMs) is costly and has significant commercial value. Consequently, preventing unauthorized appropriation of open-source LLMs and protecti…
Brain-Inspired Exploration of Functional Networks and Key Neurons in Large Language Models
Yiheng Liu, Zhengliang Liu, Zihao Wu +10
In recent years, the rapid advancement of large language models (LLMs) in natural language processing has sparked significant interest among researchers to understand their mechani…