fineweb-2 HuggingFaceFW
🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.
- 类型
- dataset
- 许可
- odc-by
- 语言
- aai
- 下载量
- 81,095
- 点赞
- 858
- 访问
- public
- 文件
- 0
标签
- 预训练语料
- 多语言语料
- 大语言模型
- 文本数据集
- 数据处理
摘要
FineWeb2是广受欢迎的FineWeb数据集的第二代迭代,为超过1000种语言提供高质量预训练语料。该数据集完全可复现,采用宽松的ODC-By 1.0许可,并通过数百次消融实验进行了广泛验证。适用于多语言大语言模型的预训练,尤其是低资源语言场景。
README
--- license: odc-by task_categories: - text-generation dataset_info: features: - name: text dtype: string - name: id dtype: string - name: dump dtype: string - name: url dtype: string - name: date dtype: string - name: file_path dtype: string - name: language dtype: string - name: language_score dt…