fineweb HuggingFaceFW

🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 datatrove library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open replication of 🦅 RefinedWeb, with a… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb.

类型
dataset
许可
odc-by
语言
en
下载量
418,229
点赞
3,211
访问
public
文件
500

标签

  • 预训练语料
  • 文本数据集
  • 大语言模型
  • 多语言语料

摘要

FineWeb 是 Hugging Face 发布的超大规模英文网络爬取文本数据集,总规模超过1万亿token,涵盖2017年至今的Common Crawl快照。它通过严格的数据清洗与去重流程产出高质量文本,广泛应用于大语言模型的预训练与持续训练。提供多种抽样规模(10BT至350BT)便于不同需求的实验与模型训练。

README

--- license: odc-by task_categories: - text-generation language: - en pretty_name: FineWeb size_categories: - n>1T configs: - config_name: default data_files: - split: train path: data/*/* - config_name: sample-10BT data_files: - split: train path: sample/10BT/* - config_name: sample-100BT data_fil…

查看完整页面 · 查看原文