activity
20242026
collaborators

12 papers

cs.CL2026

HNR-DAC: Hard-Negative Reranking and Distribution-Aligned Classification for Scientific Claim Verification

Zhenchao Wang, Xin Chen, Luoxi Zhang +2

Scientific claim verification over a cited paper requires predicting the claim--paper relation and identifying the paragraphs that justify that prediction. This setting poses two l…

cs.CL2026

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

Shuaimin Li, Liyang Fan, Zeyang Li +9

Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by exposure to benchmark data du…

cs.AI2026

EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning

Guhong Chen, Yingcheng Shi, Yongbin Li +6

Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and s…

cs.CR2026

VisPoison: An Effective Backdoor Attack Framework for Tabular Data Visualization Models

Shuaimin Li, Chen Jason Zhang, Xuanang Chen +7

Text-to-visualization (text-to-vis) models for tabular data have become essential tools in the era of big data, enabling users to generate visualizations and make data-driven decis…

cs.AI2026

Beyond Quantity: Trajectory Diversity Scaling for Code Agents

Guhong Chen, Chenghao Sun, Cheng Fu +16

As code large language models (LLMs) evolve into tool-interactive agents via the Model Context Protocol (MCP), their generalization is increasingly limited by low-quality synthetic…

cs.CL2025

Probing the Difficulty Perception Mechanism of Large Language Models

Sunbowen Lee, Qingyu Yin, Chak Tou Leong +5

Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally evaluate problem difficulty, which is an es…