wikipedia wikimedia

Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikipedia.

Typ
dataset
Lizenz
cc-by-sa-3.0
Sprache
ab
Downloads
228,662
Likes
1,369
Zugriff
public
Dateien
0

Tags

  • Mehrsprachig
  • Vortrainingskorpus
  • Textdatensatz
  • Textgenerierung
  • Maskensprachmodellierung
  • NLP
  • Große Sprachmodelle
  • Textdaten

Zusammenfassung

Dieses Datenset enthält bereinigte Wikipedia-Artikel aus fast allen Sprachversionen, abgeleitet von den offiziellen Wikipedia-Dumps. Es dient als hochwertige Trainingsgrundlage für Sprachmodellierung, Masked Language Modeling und allgemeine NLP-Vortrainingsaufgaben. Geeignet für mehrsprachige Model…

README

--- language: - ab - ace - ady - af - alt - am - ami - an - ang - anp - ar - arc - ary - arz - as - ast - atj - av - avk - awa - ay - az - azb - ba - ban - bar - bbc - bcl - be - bg - bh - bi - bjn - blk - bm - bn - bo - bpy - br - bs - bug - bxr - ca - cbk - cdo - ce - ceb - ch - chr - chy - ckb -…

查看完整页面 · 查看原文