wikipedia legacy-datasets

Wikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).

Type
dataset
License
cc-by-sa-3.0
Language
aa
Downloads
113,740
Likes
663
Access
public
Files
0

Tags

  • 多语言语料
  • 预训练语料
  • 文本数据集
  • NLP
  • 大语言模型
  • 预训练
  • 语言建模

README

--- annotations_creators: - no-annotation language_creators: - crowdsourced pretty_name: Wikipedia paperswithcode_id: null license: - cc-by-sa-3.0 - gfdl task_categories: - text-generation - fill-mask task_ids: - language-modeling - masked-language-modeling source_datasets: - original multilinguali…

查看完整页面 · 查看原文