Category: Training datasets
This category contains 4 pages.
- Common Crawl · A nonprofit's free archive of web crawls, petabytes of raw page data refreshed monthly, that serves as the raw substrate for most large language model training corpora.
- Common Pile · EleutherAI's June 2025 corpus of 8 TB of public-domain and openly licensed text from 30 sources, released with the Comma v0.1 models as an openly licensed successor to the Pile.
- ImageNet · The 14-million-image, WordNet-organized labeled dataset (2009) whose ILSVRC challenge (2010-2017) was the benchmark of the deep learning revolution, from AlexNet to ResNet.
- The Pile · EleutherAI's December 2020 corpus of 825 GiB of curated English text drawn from 22 sources, the training set behind GPT-Neo, GPT-J, GPT-NeoX-20B, and Pythia, later partly withdrawn amid the Books3 copyright dispute.