FineFineWeb m-a-p
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.
- 유형
- dataset
- 라이선스
- apache-2.0
- 언어
- en
- 다운로드
- 1,145,518
- 좋아요
- 167
- 접근
- public
- 파일
- 0
태그
- 사전학습 말뭉치
- 텍스트 데이터셋
- 데이터 필터링
- 대규모 언어 모델
- 다국어
- 대형 언어 모델
- NLP
요약
FineFineWeb은 FineWeb 원천 데이터를 77개 세부 도메인으로 정교하게 분류한 대규모 사전학습용 웹 말뭉치입니다. GPT-4 기반 URL 라벨링, Qwen2-7B-Instruct를 통한 샘플 라벨링, FastText 분류기로 세밀한 도메인 귀속을 수행해 총 4조 토큰 이상의 데이터를 제공합니다. 항공우주·의학·법률·경제 등 전문 도메인별로 특화된 LLM 사전학습과 도메인 특화 모델 개발에 적합한 데이터셋입니다.
README
--- license: apache-2.0 task_categories: - text-classification - text2text-generation - text-generation language: - en size_categories: - n>1T --- # FineFineWeb: 세밀한 도메인 웹 코퍼스에 관한 종합 연구 arXiv: 곧 공개 예정 프로젝트 페이지: 곧 공개 예정 블로그: 곧 공개 예정 ## Data Statistics | Domain (#tokens/#samples) | Iteration 1 Tokens…