7 papers
Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
Cheng Huang, Nyima Tashi, Fan Gao +19
Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research.…
BioVessel-Net and RetinaMix: Unsupervised Retinal Vessel Segmentation from OCTA Images
Cheng Huang, Weizheng Xie, Fan Gao +8
Structural changes in retinal blood vessels are critical biomarkers for the onset and progression of glaucoma and other ocular diseases. However, current vessel segmentation approa…
TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
Fan Gao, Cheng Huang, Nyima Tashi +11
To address the severe data scarcity in Tibetan, a low-resource language spoken by over six million people, we introduce TIBSTC-CoT, the large-scale, multi-domain Tibetan dataset au…
Modeling Insider Filing Delays in Financial Markets with an Interpretable XGBoost Framework
Cheng Huang, Yao Ma, Fan Gao +10
Timely disclosure of insider transactions is a cornerstone of market transparency, yet delays in filing remain widespread and challenging to monitor at scale. This study introduces…
RetrieveAll: A Multilingual Named Entity Recognition Framework with Large Language Models
Jin Zhang, Fan Gao, Linyu Li +4
The rise of large language models has led to significant performance breakthroughs in named entity recognition (NER) for high-resource languages, yet there remains substantial room…
TiSpell: A Semi-Masked Methodology for Tibetan Spelling Correction covering Multi-Level Error with Data Augmentation
Yutong Liu, Feng Xiao, Ziyue Zhang +10
Multi-level Tibetan spelling correction addresses errors at both the character and syllable levels within a unified model. Existing methods focus mainly on single-level correction…