wikipedia wikimedia

Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/wikimedia/wikipedia.

유형
dataset
라이선스
cc-by-sa-3.0
언어
ab
다운로드
228,662
좋아요
1,369
접근
public
파일
0

태그

  • 사전학습 말뭉치
  • 텍스트 데이터셋
  • 다국어
  • 대규모언어모델
  • NLP

요약

위키미디어 위키피디아 데이터셋은 위키피디아 덤프에서 추출한 모든 언어의 정제된 기사 본문을 제공합니다. 마크다운과 참고문헌 등 불필요한 섹션을 제거해 언어모델 사전학습 및 마스크 언어모델링에 적합한 고품질 텍스트 말뭉치입니다. 수백 개 언어 서브셋으로 구성되어 다국어 모델 학습과 언어모델 사전학습에 널리 활용됩니다.

README

--- language: - ab - ace - ady - af - alt - am - ami - an - ang - anp - ar - arc - ary - arz - as - ast - atj - av - avk - awa - ay - az - azb - ba - ban - bar - bbc - bcl - be - bg - bh - bi - bjn - blk - bm - bn - bo - bpy - br - bs - bug - bxr - ca - cbk - cdo - ce - ceb - ch - chr - chy - ckb -…

查看完整页面 · 查看原文