most citedAspect-Guided Multi-Level Perturbation Analysis of Large Language Models in Automated Peer Review

1 citations · 2 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2025

SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs

Zhenliang Zhang, Xinyu Hu, Xiaojun Wan

Large language models sometimes inadvertently reproduce passages that are copyrighted, exposing downstream applications to legal risk. Most existing studies for inference-time defe…

cs.CL2025

Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models

Zhenliang Zhang, Junzhe Zhang, Xinyu Hu +2

Large language models (LLMs) have achieved remarkable success in various tasks, yet they remain vulnerable to faithfulness hallucinations, where the output does not align with the…

cs.CL2025

ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs

Zhenliang Zhang, Xinyu Hu, Huixuan Zhang +2

Large language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination…

cs.CL2025

MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text

Junzhe Zhang, Huixuan Zhang, Xinyu Hu +4

Evaluation is important for multimodal generation tasks, while traditional multimodal evaluation metrics suffer from several limitations. With the rapid progress of MLLMs, there is…

cs.CL20251 cited

CFunModel: A "Funny" Language Model Capable of Chinese Humor Generation and Processing

Zhenghan Yu, Xinyu Hu, Xiaojun Wan

Humor plays a significant role in daily language communication. With the rapid development of large language models (LLMs), natural language processing has made significant strides…

cs.CL2025

Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators

Jiayi Chang, Mingqi Gao, Xinyu Hu +1

Previous research has shown that LLMs have potential in multilingual NLG evaluation tasks. However, existing research has not fully explored the differences in the evaluation capab…