4 papers
Better Datasets Start From RefineLab: Automatic Optimization for High-Quality Dataset Refinement
Xiaonan Luo, Yue Huang, Ping He +1
High-quality Question-Answer (QA) datasets are foundational for reliable Large Language Model (LLM) evaluation, yet even expert-crafted datasets exhibit persistent gaps in domain c…
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
Yue Huang, Xiangqi Wang, Xiangliang Zhang
In high-stakes scenarios-such as self-harm, legal, or medical queries-LLMs must be both trustworthy and helpful. However, these goals often conflict. We propose priority alignment,…
My Favorite Streamer is an LLM: Discovering, Bonding, and Co-Creating in AI VTuber Fandom
Jiayi Ye, Chaoran Chen, Yue Huang +3
AI VTubers, where the performer is not human but algorithmically generated, introduce a new context for fandom. While human VTubers have been substantially studied for their cultur…
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
Dawei Li, Yue Huang, Ming Li +3
Generative models such as Large Language Models, Diffusion Models, and generative adversarial networks have recently revolutionized the creation of synthetic data, offering scalabl…