2 papers
cs.CL2024
REInstruct: Building Instruction Data from Unlabeled Corpus
Shu Chen, Xinyan Guan, Yaojie Lu +3
Manually annotating instruction data for large language models is difficult, costly, and hard to scale. Meanwhile, current automatic annotation methods typically rely on distilling…
cs.CL2023
DLUE: Benchmarking Document Language Understanding
Ruoxi Xu, Hongyu Lin, Xinyan Guan +3
Understanding documents is central to many real-world tasks but remains a challenging topic. Unfortunately, there is no well-established consensus on how to comprehensively evaluat…