collaborators

6 papers

cs.CL2026

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

Qingjie Zhang, Xingzhang Ren, Zixuan Chen +6

Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific…

cs.CL2026

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

Qingjie Zhang, Ziqi Tang, Jie Zhang +7

Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three challenges: (1) their web-scale size makes ful…

eess.SP2026

The Note-Chord-Voice Framework: Structured Source Separation and Causal Inference for EV Charging Data

Jiajie Chen, Jinfeng Li

Real-world EV charging data exhibit three interlocking pathologies: hardware fragmentation (network timeouts and billing resets split sessions), physical violations (independent en…

cs.CL2026

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

Junyu Lin, Meizhen Liu, Xiufeng Huang +12

As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained filtering and support fine-grained,…

cs.CL2024

fairBERTs: Erasing Sensitive Information Through Semantic and Fairness-aware Perturbations

Jinfeng Li, Yuefeng Chen, Xiangyu Liu +3

Pre-trained language models (PLMs) have revolutionized both the natural language processing research and applications. However, stereotypical biases (e.g., gender and racial discri…

cs.CR2024

S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Xiaohan Yuan, Jinfeng Li, Dongxia Wang +9

Generative large language models (LLMs) have revolutionized natural language processing with their transformative and emergent capabilities. However, recent evidence indicates that…