collaborators

5 papers

cs.CL2026

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models

Kaixin Lan, Mu You, Tao Fang +3

Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and raises ethical, legal, and fairnes…

cs.CL2026

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

Xiaoqi He, Kaixin Lan, Mu You +3

Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs) in ancient Chinese texts. Th…

cs.CL2026

CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models

Yike Sun, Mingkun Xu, Mu You +5

In recent years, the black-box nature of deep learning models has limited their application in high-stakes domains such as medical diagnosis and finance, where interpretability is…

cs.CL2024

LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning

Tao Fang, Derek F. Wong, Lusheng Zhang +5

While large-scale language models (LLMs) have demonstrated remarkable capabilities in specific natural language processing (NLP) tasks, they may still lack proficiency compared to…

cs.CL2024

FOCUS: Forging Originality through Contrastive Use in Self-Plagiarism for Language Models

Kaixin Lan, Tao Fang, Derek F. Wong +3

Pre-trained Language Models (PLMs) have shown impressive results in various Natural Language Generation (NLG) tasks, such as powering chatbots and generating stories. However, an e…