finepdfs HuggingFaceFW

Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs.

種別
dataset
ライセンス
odc-by
言語
aai
ダウンロード
45,951
いいね
914
アクセス
public
ファイル
0

タグ

  • 事前学習コーパス
  • 多言語コーパス
  • テキストデータ
  • 大規模データセット
  • NLP
  • 大言語モデル
  • データセット

概要

FinePDFsは、PDFのみをソースとする最大の公開コーパスで、約3兆トークン・4億7500万文書・1733言語を収録しています。Webページ枯渇に伴い、抽出コストが高くこれまで敬遠されていたPDFを大規模に解放することを目的とし、多言語の大規模言語モデル事前学習用データとして高品質なテキスト生成・言語モデリングに利用できます。

README

--- license: odc-by task_categories: - text-generation pretty_name: 📄 FinePDFs language: - aai - aak - aau - aaz - aba - abi - abk - abn - abq - abs - abt - abx - aby - abz - aca - acd - ace - acf - ach - acm - acn - acr - acu - ada - ade - adh - adi - adj - adl - ady - adz - aeb - aer - aeu - aey…

查看完整页面 · 查看原文