5 papers · 1 filter
From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models
Jiaxu Zuo, Mu You, Kaixin Lan +5
Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly…
MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models
Kaixin Lan, Mu You, Tao Fang +3
Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and raises ethical, legal, and fairnes…
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
Xiaoqi He, Kaixin Lan, Mu You +3
Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs) in ancient Chinese texts. Th…
Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
Yanming Sun, Runzhe Zhan, Chi Seng Cheang +7
\textbf{RE}trieval-\textbf{A}ugmented \textbf{L}LM-based \textbf{M}achine \textbf{T}ranslation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but…
FOCUS: Forging Originality through Contrastive Use in Self-Plagiarism for Language Models
Kaixin Lan, Tao Fang, Derek F. Wong +3
Pre-trained Language Models (PLMs) have shown impressive results in various Natural Language Generation (NLG) tasks, such as powering chatbots and generating stories. However, an e…