Publications (13)
HLB: Benchmarking LLMs' Humanlikeness in Language Use
Xufeng Duan, Bei Xiao, Xuemei Tang +1
As synthetic data becomes increasingly prevalent in training language models, particularly through generated dialogue, concerns have emerged that these models may deviate from auth…
That Slepen Al the Nyght with Open Ye! Cross-era Sequence Segmentation with Switch-memory
Xuemei Tang, Qi Su, Jun Wang
The evolution of language follows the rule of gradual change. Grammar, vocabulary, and lexical semantic shifts take place over time, resulting in a diachronic linguistic gap. As su…
CHisIEC: An Information Extraction Corpus for Ancient Chinese History
Xuemei Tang, Zekun Deng, Qi Su +2
Natural Language Processing (NLP) plays a pivotal role in the realm of Digital Humanities (DH) and serves as the cornerstone for advancing the structural analysis of historical and…
From Pretraining to Privacy: Federated Ultrasound Foundation Model with Self-Supervised Learning
Yuncheng Jiang, Chun-Mei Feng, Jinke Ren +15
Ultrasound imaging is widely used in clinical diagnosis due to its non-invasive nature and real-time capabilities. However, traditional ultrasound diagnostics relies heavily on phy…
Small Language Models as Effective Guides for Large Language Models in Chinese Relation Extraction
Xuemei Tang, Jun Wang
Recently, large language models (LLMs) have been successful in relational extraction (RE) tasks, especially in the few-shot learning. An important problem in the field of RE is lon…
Towards a Benchmark for Colorectal Cancer Segmentation in Endorectal Ultrasound Videos: Dataset and Model Development
Yuncheng Jiang, Yiwen Hu, Zixun Zhang +7
Endorectal ultrasound (ERUS) is an important imaging modality that provides high reliability for diagnosing the depth and boundary of invasion in colorectal cancer. However, the la…