1 paper
Stephan Oepen, Nikolay Arefev, Mikko Aulamo +29
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely th…