20 papers
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generation
Zeli Su, Ziyin Zhang, Zewei Pan +8
Low-resource target-language generation is often limited by scarce parallel data, while high-resource source-language monolingual data is abundant but difficult to use with standar…
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
Guixian Xu, Yide Liang, Zeli Su +5
Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastruct…
TingIS: Real-time Risk Event Discovery from Noisy Customer Incidents at Enterprise Scale
Jun Wang, Ziyin Zhang, Rui Wang +2
Real-time detection and mitigation of technical anomalies are critical for large-scale cloud-native services, where even minutes of downtime can result in massive financial losses…
ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World
Ziyin Zhang, Zihan Liao, Hang Yu +2
The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitive computational costs, a narro…
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
Zeli Su, Ziyin Zhang, Zhou Liu +7
Extending large language models (LLMs) to low-resource languages often incurs an "alignment tax": improvements in the target language come at the cost of catastrophic forgetting in…
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
Siqiao Xue, Zihan Liao, Jin Qin +4
Code search has usually been evaluated as first-stage retrieval, even though production systems rely on broader pipelines with reranking and developer-style queries. Existing bench…