318 citations · 619 across the 60 of their papers we have counts for
1 paper · 2 filters
Jeffrey Li, Josh Gardner, Doug Kang +10
One of the first pre-processing steps for constructing web-scale LLM pretraining datasets involves extracting text from HTML. Despite the immense diversity of web content, existing…