24 citations · 44 across the 7 of their papers we have counts for
7 papers
How to Train Data-Efficient LLMs
Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang +6
The training of large language models (LLMs) is expensive. In this paper, we study data-efficient approaches for pre-training LLMs, i.e., techniques that aim to optimize the Pareto…
Farzi Data: Autoregressive Data Distillation
Noveen Sachdeva, Zexue He, Wang-Cheng Kang +3
We study data distillation for auto-regressive machine learning tasks, where the input and output have a strict left-to-right causal structure. More specifically, we propose Farzi,…
Do LLMs Understand User Preferences? Evaluating LLMs On User Rating Prediction
Wang-Cheng Kang, Jianmo Ni, Nikhil Mehta +4
Large Language Models (LLMs) have demonstrated exceptional capabilities in generalizing to new tasks in a zero-shot or few-shot manner. However, the extent to which LLMs can compre…
WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset
Andrea Burns, Krishna Srinivasan, Joshua Ainslie +5
Webpages have been a rich resource for language and vision-language tasks. Yet only pieces of webpages are kept: image-caption pairs, long text articles, or raw HTML, never all in…
Knowledge-aware Neural Collective Matrix Factorization for Cross-domain Recommendation
Li Zhang, Yan Ge, Jun Ma +2
Cross-domain recommendation (CDR) can help customers find more satisfying items in different domains. Existing CDR models mainly use common users or mapping functions as bridges be…
Large Dual Encoders Are Generalizable Retrievers
Jianmo Ni, Chen Qu, Jing Lu +8
It has been shown that dual encoders trained on one domain often fail to generalize to other domains for retrieval tasks. One widespread belief is that the bottleneck layer of a du…