fineweb-2 HuggingFaceFW

🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.

Tipo
dataset
Licencia
odc-by
Lenguaje
aai
Descargas
81,095
Me gusta
858
Acceso
public
Archivos
0

Etiquetas

  • preentrenamiento de LLM
  • multilingüe
  • corpus de texto
  • texto web
  • preentrenamiento
  • NLP
  • LLM

Resumen

FineWeb2 es la segunda iteración del popular dataset FineWeb, que aporta datos de preentrenamiento de alta calidad para más de 1000 idiomas. Es totalmente reproducible, publicado bajo licencia permisiva ODC-By 1.0 y validado mediante cientos de experimentos de ablación. Sirve para entrenar y preent…

README

--- license: odc-by task_categories: - text-generation dataset_info: features: - name: text dtype: string - name: id dtype: string - name: dump dtype: string - name: url dtype: string - name: date dtype: string - name: file_path dtype: string - name: language dtype: string - name: language_score dt…

查看完整页面 · 查看原文