smollm-corpus HuggingFaceTB

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.

유형
dataset
라이선스
odc-by
언어
en
다운로드
46,086
좋아요
481
접근
public
파일
0

태그

  • 사전학습 말뭉치
  • 합성 데이터
  • 교육 데이터셋
  • 코드 데이터
  • LLM
  • 데이터 필터링
  • 딥러닝

요약

SmolLM-Corpus는 소형 언어 모델(SmolLM) 훈련을 위해 설계된 선별된 고품질 교육 데이터 및 합성 데이터 모음입니다. 대규모 합성 데이터셋인 Cosmopedia v2(Mixtral-8x7B로 생성한 3900만 개 이상의 교과서·블로그·이야기), 교육용 웹 데이터를 전용 분류기로 필터링한 FineWeb-Edu-Dedup, 교육적 코드 점수로 선별된 python-edu 서브셋으로 구성됩니다. 소형 모델 사전학습 성능 향상에 활용되는 다목적 교육용 멀티모달 코퍼스입니다.

README

--- license: odc-by dataset_info: - config_name: cosmopedia-v2 features: - name: prompt dtype: string - name: text dtype: string - name: token_length dtype: int64 - name: audience dtype: string - name: format dtype: string - name: seed_data dtype: string splits: - name: train num_bytes: 21250364074…

查看完整页面 · 查看原文