fineweb-2 HuggingFaceFW
🥂 FineWeb2 A sparkling update with 1000s of languages What is it? This is the second iteration of the popular 🍷 FineWeb dataset, bringing high quality pretraining data to over 1000 🗣️ languages. The 🥂 FineWeb2 dataset is fully reproducible, available under the permissive ODC-By 1.0 license and extensively validated through hundreds of ablation experiments. In particular, on the set of 9 diverse languages we used to guide our processing decisions, 🥂… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-2.
- 유형
- dataset
- 라이선스
- odc-by
- 언어
- aai
- 다운로드
- 81,095
- 좋아요
- 858
- 접근
- public
- 파일
- 0
태그
- 사전학습 말뭉치
- 다국어
- 텍스트 데이터셋
- 데이터 필터링
- 대규모 언어 모델
- NLP
요약
FineWeb2는 인기 있는 FineWeb 데이터셋의 두 번째 버전으로, 1000개 이상의 언어에 걸친 고품질 사전학습(pre-training) 텍스트 데이터를 제공합니다. 완전히 재현 가능하며 ODC-By 1.0 허용 라이선스로 공개되었고, 수백 건의 제거(ablation) 실험으로 검증되었습니다. 다국어 대규모 언어 모델 사전학습에 최적화된 말뭉치로, 저자원 언어 포함 다양한 언어 커버리지를 지원합니다.
README
--- license: odc-by task_categories: - text-generation dataset_info: features: - name: text dtype: string - name: id dtype: string - name: dump dtype: string - name: url dtype: string - name: date dtype: string - name: file_path dtype: string - name: language dtype: string - name: language_score dt…