collaborators

5 papers

cs.CV2026

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

Guixian Xu, Yide Liang, Zeli Su +5

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastruct…

cs.CL2026

Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax

Zeli Su, Ziyin Zhang, Zhou Liu +7

Extending large language models (LLMs) to low-resource languages often incurs an "alignment tax": improvements in the target language come at the cost of catastrophic forgetting in…

cs.LG2025

SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression

Zeli Su, Ziyin Zhang, Wenzheng Zhang +3

Transformer encoders are widely deployed in large-scale web services for natural language understanding tasks such as text classification, semantic retrieval, and content ranking.…

cs.CL2025

CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

Guixian Xu, Zeli Su, Ziyin Zhang +4

Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international st…

cs.CL2025

Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages

Zeli Su, Ziyin Zhang, Guixian Xu +4

While multilingual language models like XLM-R have advanced multilingualism in NLP, they still perform poorly in extremely low-resource languages. This situation is exacerbated by…