fineweb-2 HuggingFaceFW
🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.
- Tipo
- dataset
- Licencia
- odc-by
- Lenguaje
- aai
- Descargas
- 81,095
- Me gusta
- 858
- Acceso
- public
- Archivos
- 0
Etiquetas
- preentrenamiento de LLM
- multilingüe
- corpus de texto
- texto web
- preentrenamiento
- NLP
- LLM
Resumen
FineWeb2 es la segunda iteración del popular dataset FineWeb, que aporta datos de preentrenamiento de alta calidad para más de 1000 idiomas. Es totalmente reproducible, publicado bajo licencia permisiva ODC-By 1.0 y validado mediante cientos de experimentos de ablación. Sirve para entrenar y preent…
README
--- license: odc-by task_categories: - text-generation dataset_info: features: - name: text dtype: string - name: id dtype: string - name: dump dtype: string - name: url dtype: string - name: date dtype: string - name: file_path dtype: string - name: language dtype: string - name: language_score dt…