collaborators

5 papers

cs.CL2025

DAST: Difficulty-Aware Self-Training on Large Language Models

Boyang Xue, Qi Zhu, Hongru Wang +8

Present Large Language Models (LLM) self-training methods always under-sample on challenging queries, leading to inadequate learning on difficult problems which limits LLMs' abilit…

cs.CL2025

Benchmarking Large Language Models on Multiple Tasks in Bioinformatics NLP with Prompting

Jiyue Jiang, Pengan Chen, Jiuming Wang +13

Large language models (LLMs) have become important tools in solving biological problems, offering improvements in accuracy and adaptability over conventional methods. Several bench…

cs.CL2025

Developing and Utilizing a Large-Scale Cantonese Dataset for Multi-Tasking in Large Language Models

Jiyue Jiang, Alfred Kar Yin Truong, Yanyu Chen +7

High-quality data resources play a crucial role in learning large language models (LLMs), particularly for low-resource languages like Cantonese. Despite having more than 85 millio…

cs.LG2025

TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning

Sheng Wang, Pengan Chen, Jingqi Zhou +7

Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large languag…

cs.CL2024

UAlign: Leveraging Uncertainty Estimations for Factuality Alignment on Large Language Models

Boyang Xue, Fei Mi, Qi Zhu +6

Despite demonstrating impressive capabilities, Large Language Models (LLMs) still often struggle to accurately express the factual knowledge they possess, especially in cases where…