FineFineWeb m-a-p

FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.

유형
dataset
라이선스
apache-2.0
언어
en
다운로드
1,145,518
좋아요
167
접근
public
파일
0

태그

  • 사전학습 말뭉치
  • 텍스트 데이터셋
  • 데이터 필터링
  • 대규모 언어 모델
  • 다국어
  • 대형 언어 모델
  • NLP

요약

FineFineWeb은 FineWeb 원천 데이터를 77개 세부 도메인으로 정교하게 분류한 대규모 사전학습용 웹 말뭉치입니다. GPT-4 기반 URL 라벨링, Qwen2-7B-Instruct를 통한 샘플 라벨링, FastText 분류기로 세밀한 도메인 귀속을 수행해 총 4조 토큰 이상의 데이터를 제공합니다. 항공우주·의학·법률·경제 등 전문 도메인별로 특화된 LLM 사전학습과 도메인 특화 모델 개발에 적합한 데이터셋입니다.

README

--- license: apache-2.0 task_categories: - text-classification - text2text-generation - text-generation language: - en size_categories: - n>1T --- # FineFineWeb: 세밀한 도메인 웹 코퍼스에 관한 종합 연구 arXiv: 곧 공개 예정 프로젝트 페이지: 곧 공개 예정 블로그: 곧 공개 예정 ## Data Statistics | Domain (#tokens/#samples) | Iteration 1 Tokens…

查看完整页面 · 查看原文