finepdfs HuggingFaceFW
Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.
- Tipo
- dataset
- Licencia
- odc-by
- Lenguaje
- aai
- Descargas
- 45,951
- Me gusta
- 914
- Acceso
- public
- Archivos
- 0
Etiquetas
- preentrenamiento de LLM
- corpus de texto
- multilingüe
- NLP
- texto
- preentrenamiento
Resumen
FinePDFs es el corpus público más grande extraído exclusivamente de PDFs, con aproximadamente 3 billones de tokens en 475 millones de documentos que abarcan 1733 idiomas. Su propósito es liberar una nueva fuente de datos de alta calidad cuando se agotan las páginas web disponibles. Es ideal para pr…
README
--- license: odc-by task_categories: - text-generation pretty_name: 📄 FinePDFs language: - aai - aak - aau - aaz - aba - abi - abk - abn - abq - abs - abt - abx - aby - abz - aca - acd - ace - acf - ach - acm - acn - acr - acu - ada - ade - adh - adi - adj - adl - ady - adz - aeb - aer - aeu - aey…