c4 allenai
C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.
- 유형
- dataset
- 라이선스
- odc-by
- 언어
- af
- 다운로드
- 1,430,934
- 좋아요
- 630
- 접근
- public
- 파일
- 0
태그
- 사전학습 말뭉치
- 다국어
- 텍스트 데이터셋
- 대규모 언어 모델
- NLP
- 마스크 언어모델
요약
C4는 Common Crawl의 웹 크롤링 말뭉치를 대규모로 정제해 만든 사전학습용 텍스트 데이터셋입니다. 영어(en)를 비롯해 다국어(mC4) 등 5가지 변형을 제공하며, LLM 사전학습 및 마스크 언어모델 학습에 널리 사용됩니다. 수백 기가바이트에서 테라바이트에 이르는 막대한 규모의 언어 모델 파운데이션 데이터로 활용됩니다.
README
--- pretty_name: C4 annotations_creators: - no-annotation language_creators: - found language: - af - am - ar - az - be - bg - bn - ca - ceb - co - cs - cy - da - de - el - en - eo - es - et - eu - fa - fi - fil - fr - fy - ga - gd - gl - gu - ha - haw - he - hi - hmn - ht - hu - hy - id - ig - is …