finepdfs HuggingFaceFW
Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.
- 유형
- dataset
- 라이선스
- odc-by
- 언어
- aai
- 다운로드
- 45,951
- 좋아요
- 914
- 접근
- public
- 파일
- 0
태그
- 사전학습 말뭉치
- 다국어
- 텍스트 생성
- 데이터셋
- 대규모 언어 모델
- NLP
요약
FinePDFs는 PDF 소스에서만 추출된 가장 큰 공개 말뭉치로, 1733개 언어에 걸친 약 4억 7,500만 개 문서와 3조 개에 달하는 토큰으로 구성된 사전학습용 데이터셋입니다. 웹 페이지 데이터 부족 문제를 해결하기 위해 기존에 높은 추출 비용과 복잡성으로 기피되던 PDF라는 데이터 소스를 대규모로 정제·활용합니다. 대규모 언어모델의 사전학습 및 다국어 모델 학습에 주로 활용할 수 있습니다.
README
--- license: odc-by task_categories: - text-generation pretty_name: 📄 FinePDFs language: - aai - aak - aau - aaz - aba - abi - abk - abn - abq - abs - abt - abx - aby - abz - aca - acd - ace - acf - ach - acm - acn - acr - acu - ada - ade - adh - adi - adj - adl - ady - adz - aeb - aer - aeu - aey…