pile EleutherAI
The Pile is a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets combined together.
- 유형
- dataset
- 라이선스
- other
- 언어
- en
- 다운로드
- 2,877
- 좋아요
- 502
- 접근
- public
- 파일
- 0
태그
- 사전학습 말뭉치
- 텍스트 데이터셋
- 대규모 언어 모델
- 다국어
- 벤치마크
- NLP
요약
The Pile은 EleutherAI가 구축한 825GiB 용량의 다양하고 고품질인 오픈소스 언어모델 학습용 데이터셋으로, 22개의 하위 데이터셋을 결합해 구성되었습니다. 의학 논문, 법률 문서, 뉴스, 코드, 대화 등 다양한 도메인의 영어 텍스트를 포함하여 언어모델 사전학습에 폭넓게 활용됩니다. 대규모 언어모델의 사전학습 말뭉치로 사용되며, GPT-Neo 등 다양한 오픈소스 LLM 학습에 기반이 된 대표적인 벤치마크 데이터셋입니다.
README
--- annotations_creators: - no-annotation language_creators: - found language: - en license: other multilinguality: - monolingual pretty_name: the Pile size_categories: - 100B<n<1T source_datasets: - original task_categories: - text-generation - fill-mask task_ids: - language-modeling - masked-lang…