collaborators

15 papers

cs.CL2026

Do Large Language Models Perform Well on Comprehending Poetic Logic in Modern Chinese Poetry?

Tian Lan, Shanshan Wang, Zehua Duo +4

Large Language Models (LLMs) have achieved significant progress across a wide range of natural language processing (NLP) tasks, yet their ability to understand literary texts, part…

cs.CL2026

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?

Shanshan Wang, Derek F. Wong, Jingming Yao +1

Traditional automatic evaluation methods have been shown to be unsuitable for modern Chinese poetry because of the distinct nature of this literary genre. Human evaluation remains…

cs.CL2026

Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity

Yuanhe Zhao, Tianyu Zhang, Huafei Xing +3

Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios risks privacy leakage via malicious promp…

cs.CL2026

From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

Jiaxu Zuo, Mu You, Kaixin Lan +5

Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly…

cs.CL2026

G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment

Fengying Ye, Yanming Sun, Runzhe Zhan +3

Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a…

cs.CL2026

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models

Kaixin Lan, Mu You, Tao Fang +3

Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and raises ethical, legal, and fairnes…