collaborators

11 papers

cs.AI2026

AI Evaluation Should Require Standardized Item-Level Data Releases

Han Jiang, Susu Zhang, Dongyao Zhu +6

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…

cs.IR2026

MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints

Kangyang Luo, Shuzheng Si, Yuzhuo Bai +8

In the era of large language models (LLMs), supervised neural methods remain the state-of-the-art (SOTA) for Coreference Resolution. Yet, their full potential is underexplored, par…

cs.CL2026

ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement

Kangyang Luo, Yuzhuo Bai, Shuzheng Si +9

Coreference Resolution (CR) is a critical task in Natural Language Processing (NLP). Current research faces a key dilemma: whether to further explore the potential of supervised ne…

cs.CL2026

FaithLens: Detecting and Explaining Faithfulness Hallucination

Shuzheng Si, Qingyi Wang, Haozhe Zhao +8

Recognizing whether outputs from large language models (LLMs) contain faithfulness hallucination is crucial for real-world applications, e.g., retrieval-augmented generation and su…

cs.CL2026

InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs

Yuzhuo Bai, Shuzheng Si, Kangyang Luo +5

Large language models (LLMs) often hallucinate, yet most existing fact-checking methods treat factuality evaluation as a binary classification problem, offering limited interpretab…

cs.CL2025

IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization

Yuzhuo Bai, Shitong Duan, Muhua Huang +7

Trained on various human-authored corpora, Large Language Models (LLMs) have demonstrated a certain capability of reflecting specific human-like traits (e.g., personality or values…