7 papers · 1 filter
TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs
Jin Zhang, Linyu Li, Weili Jiang +7
Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions…
TFD: A Comprehensive Structured Tibetan Foundation Dataset for Low-Resource Language Processing and Large-Scale Modeling
Cheng Huang, Fan Gao, Nyima Tashi +5
Large Language Models (LLMs) have achieved remarkable success in high-resource languages, yet progress in Tibetan remains severely constrained. While recent efforts have begun to a…
TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ã-Tsang, Amdo and Kham Speech Dataset Generation
Yutong Liu, Ziyue Zhang, Ban Ma-bao +7
Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (Ã-Tsang, Amdo, and Kham), limiting progress in speech modeling. To addre…
TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
Fan Gao, Cheng Huang, Nyima Tashi +11
To address the severe data scarcity in Tibetan, a low-resource language spoken by over six million people, we introduce TIBSTC-CoT, the large-scale, multi-domain Tibetan dataset au…
Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
Cheng Huang, Nyima Tashi, Fan Gao +19
Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research.…
TLUE: A Tibetan Language Understanding Evaluation Benchmark
Fan Gao, Cheng Huang, Nyima Tashi +9
Large language models have made tremendous progress in recent years, but low-resource languages, like Tibetan, remain significantly underrepresented in their evaluation. Despite Ti…