178 citations · 560 across the 43 of their papers we have counts for
19 papers · 1 filter
DataComp-LM: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis +56
We introduce DataComp for Language Models (DCLM), a testbed for controlled dataset experiments with the goal of improving language models. As part of DCLM, we provide a standardize…
Taming Small-sample Bias in Low-budget Active Learning
Linxin Song, Jieyu Zhang, Xiaotian Lu +1
Active learning (AL) aims to minimize the annotation cost by only querying a few informative examples for each model training stage. However, training a model on a few queried exam…
On the Trade-off of Intra-/Inter-class Diversity for Supervised Pre-training
Jieyu Zhang, Bohan Wang, Zhengyu Hu +2
Pre-training datasets are critical for building state-of-the-art machine learning models, motivating rigorous study on their impact on downstream tasks. In this work, we study the…
Label-Efficient Interactive Time-Series Anomaly Detection
Hong Guo, Yujing Wang, Jieyu Zhang +5
Time-series anomaly detection is an important task and has been widely applied in the industry. Since manual data annotation is expensive and inefficient, most applications adopt u…
Single-Pass Contrastive Learning Can Work for Both Homophilic and Heterophilic Graph
Haonan Wang, Jieyu Zhang, Qi Zhu +3
Existing graph contrastive learning (GCL) techniques typically require two forward passes for a single instance to construct the contrastive loss, which is effective for capturing…
Leveraging Instance Features for Label Aggregation in Programmatic Weak Supervision
Jieyu Zhang, Linxin Song, Alexander Ratner
Programmatic Weak Supervision (PWS) has emerged as a widespread paradigm to synthesize training labels efficiently. The core component of PWS is the label model, which infers true…