22 papers
Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?
Shanshan Wang, Derek F. Wong, Jingming Yao +1
Traditional automatic evaluation methods have been shown to be unsuitable for modern Chinese poetry because of the distinct nature of this literary genre. Human evaluation remains…
From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models
Jiaxu Zuo, Mu You, Kaixin Lan +5
Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly…
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
Fengying Ye, Yanming Sun, Runzhe Zhan +3
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mappings unreliable. We present G-IdiomAlign, a…
Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
Fengying Ye, Shanshan Wang, Lidia S. Chao +1
Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such behavioral success reveals about metaphor p…
MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models
Kaixin Lan, Mu You, Tao Fang +3
Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and raises ethical, legal, and fairnes…
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
Xiaoqi He, Kaixin Lan, Mu You +3
Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs) in ancient Chinese texts. Th…