collaborators

6 papers

cs.CL2026

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

Shuaimin Li, Liyang Fan, Zeyang Li +9

Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by exposure to benchmark data du…

cs.RO2026

AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control

Yutian Cheng, Xiaojian Ma, Xianhao Wang +6

Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational ov…

cs.CL2026

Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debates

Shuaimin Li, Liyang Fan, Yufang Lin +5

Existing paper review methods often rely on superficial manuscript features or directly on large language models (LLMs), which are prone to hallucinations, biased scoring, and limi…

cs.AI2026

CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs

Siyi Li, Jiajun Shi, Shiwen Ni +9

Large Reasoning Models (LRMs) have demonstrated strong performance by producing extended Chain-of-Thought (CoT) traces before answering. However, this paradigm often induces over-r…

cs.CL2025

A Survey on Large Language Model Benchmarks

Shiwen Ni, Guhong Chen, Shuaimin Li +11

In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in incre…

cs.CL2025

xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking

Sunbowen Lee, Shiwen Ni, Chi Wei +7

Safety alignment mechanism are essential for preventing large language models (LLMs) from generating harmful information or unethical content. However, cleverly crafted prompts can…