A growing and diverse dataset of text for AI to graze on and learn new information. Just like a pasture in the wild, it is a combination of sources. All the data is in Arrow format so it is easy to randomly access and stream.