finepdfs HuggingFaceFW
Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.
- 種別
- dataset
- ライセンス
- odc-by
- 言語
- aai
- ダウンロード
- 45,951
- いいね
- 914
- アクセス
- public
- ファイル
- 0
タグ
- 事前学習コーパス
- 多言語コーパス
- テキストデータ
- 大規模データセット
- NLP
- 大言語モデル
- データセット
概要
FinePDFsは、PDFのみをソースとする最大の公開コーパスで、約3兆トークン・4億7500万文書・1733言語を収録しています。Webページ枯渇に伴い、抽出コストが高くこれまで敬遠されていたPDFを大規模に解放することを目的とし、多言語の大規模言語モデル事前学習用データとして高品質なテキスト生成・言語モデリングに利用できます。
README
--- license: odc-by task_categories: - text-generation pretty_name: 📄 FinePDFs language: - aai - aak - aau - aaz - aba - abi - abk - abn - abq - abs - abt - abx - aby - abz - aca - acd - ace - acf - ach - acm - acn - acr - acu - ada - ade - adh - adi - adj - adl - ady - adz - aeb - aer - aeu - aey…