18 citations · 108 across the 32 of their papers we have counts for
19 papers · 1 filter
BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training
Jiaxing Wang, Deping Xiang, Jin Xu +9
As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive lear…
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
Zijie Liu, Jie Peng, Jinhao Duan +7
Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets. However, SM…
Chasing Fairness in Graphs: A GNN Architecture Perspective
Zhimeng Jiang, Xiaotian Han, Chao Fan +4
There has been significant progress in improving the performance of graph neural networks (GNNs) through enhancements in graph data, model architecture design, and training strateg…
TVE: Learning Meta-attribution for Transferable Vision Explainer
Guanchu Wang, Yu-Neng Chuang, Fan Yang +8
Explainable machine learning significantly improves the transparency of deep neural networks. However, existing work is constrained to explaining the behavior of individual model p…
Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through Honeypots
Ruixiang Tang, Jiayi Yuan, Yiming Li +3
In the field of natural language processing, the prevalent approach involves fine-tuning pretrained language models (PLMs) using local samples. Recent research has exposed the susc…
Efficient GNN Explanation via Learning Removal-based Attribution
Yao Rong, Guanchu Wang, Qizhang Feng +4
As Graph Neural Networks (GNNs) have been widely used in real-world applications, model explanations are required not only by users but also by legal regulations. However, simultan…