collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models

Jiaxu Zuo, Mu You, Kaixin Lan +5

Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms underlying LLM-based scoring remain poorly…

cs.CL2026

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models

Kaixin Lan, Mu You, Tao Fang +3

Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and raises ethical, legal, and fairnes…

cs.CL2026

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

Xiaoqi He, Kaixin Lan, Mu You +3

Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs) in ancient Chinese texts. Th…

cs.CL2025

Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation

Yanming Sun, Runzhe Zhan, Chi Seng Cheang +7

\textbf{RE}trieval-\textbf{A}ugmented \textbf{L}LM-based \textbf{M}achine \textbf{T}ranslation (REAL-MT) shows promise for knowledge-intensive tasks like idiomatic translation, but…

cs.CL2024

FOCUS: Forging Originality through Contrastive Use in Self-Plagiarism for Language Models

Kaixin Lan, Tao Fang, Derek F. Wong +3

Pre-trained Language Models (PLMs) have shown impressive results in various Natural Language Generation (NLG) tasks, such as powering chatbots and generating stories. However, an e…