9 papers
SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models
Shuaimin Li, Liyang Fan, Zeyang Li +9
Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by exposure to benchmark data du…
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Haoxiang Guan, Jiyan He, Liyang Fan +10
Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simulations. Traditional agent-based models (…
Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates
Shuaimin Li, Liyang Fan, Yufang Lin +5
Existing paper review methods often rely on superficial manuscript features or directly on large language models (LLMs), which are prone to hallucinations, biased scoring, and limi…
DoPE: Denoising Rotary Position Embedding
Jing Xiong, Liyang Fan, Hui Shen +4
Positional encoding is essential for large language models (LLMs) to represent sequence order, yet recent studies show that Rotary Position Embedding (RoPE) can induce massive acti…
IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property
Qiyao Wang, Guhong Chen, Hongbo Wang +20
Intellectual Property (IP) is a highly specialized domain that integrates technical and legal knowledge, making it inherently complex and knowledge-intensive. Recent advancements i…
A Survey on Large Language Model Benchmarks
Shiwen Ni, Guhong Chen, Shuaimin Li +11
In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in incre…