fineweb-2 HuggingFaceFW

🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.

유형
dataset
라이선스
odc-by
언어
aai
다운로드
81,095
좋아요
858
접근
public
파일
0

태그

  • 사전학습 말뭉치
  • 다국어
  • 텍스트 데이터셋
  • 데이터 필터링
  • 대규모 언어 모델
  • NLP

요약

FineWeb2는 인기 있는 FineWeb 데이터셋의 두 번째 버전으로, 1000개 이상의 언어에 걸친 고품질 사전학습(pre-training) 텍스트 데이터를 제공합니다. 완전히 재현 가능하며 ODC-By 1.0 허용 라이선스로 공개되었고, 수백 건의 제거(ablation) 실험으로 검증되었습니다. 다국어 대규모 언어 모델 사전학습에 최적화된 말뭉치로, 저자원 언어 포함 다양한 언어 커버리지를 지원합니다.

README

--- license: odc-by task_categories: - text-generation dataset_info: features: - name: text dtype: string - name: id dtype: string - name: dump dtype: string - name: url dtype: string - name: date dtype: string - name: file_path dtype: string - name: language dtype: string - name: language_score dt…

查看完整页面 · 查看原文