wikitext Salesforce

Dataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.

유형
dataset
라이선스
cc-by-sa-3.0
언어
en
다운로드
1,498,204
좋아요
758
접근
public
파일
0

태그

  • 사전학습 말뭉치
  • 텍스트 생성
  • 마스크 언어모델
  • NLP
  • 벤치마크
  • 데이터셋
  • 다국어

요약

위키텍스트(WikiText)는 위키백과의 검증된 우수/추천 문서에서 추출한 1억 개 이상 토큰으로 구성된 영어 언어모델링 데이터셋입니다. 기존 Penn Treebank(PTB)보다 훨씬 큰 어휘와 원문의 대소문자·구두점을 보존하여 장기 의존성(long-term dependency) 학습에 적합하며, 언어모델 성능 평가의 표준 벤치마크로 널리 사용됩니다. 원시(문자 단위)용 raw 버전과 단어 단위용 non-raw 버전, 그리고 위키텍스트-2와 위키텍스트-103 두 가지 크기 변형을 제공합니다. 사전학습 말뭉치로도 활용되지만 주로 …

README

--- annotations_creators: - no-annotation language_creators: - crowdsourced language: - en license: - cc-by-sa-3.0 - gfdl multilinguality: - monolingual size_categories: - 1M<n<10M source_datasets: - original task_categories: - text-generation - fill-mask task_ids: - language-modeling - masked-lang…

查看完整页面 · 查看原文