MINT-1T-HTML mlfoundations

🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.

Typ
dataset
Lizenz
cc-by-4.0
Sprache
en
Downloads
186,938
Likes
97
Zugriff
public
Dateien
0

Tags

  • 多模态数据集
  • Vortrainingskorpus
  • 多模态
  • 预训练语料库
  • Bilddatensatz
  • Textdatensatz
  • 英语模型
  • Deduplizierung

Zusammenfassung

MINT-1T ist ein hochskalierter multimodaler, verschachtelter Datensatz mit 1 Billion Text-Tokens und 3,4 Milliarden Bildern – eine 10-fache Skalierung gegenüber bestehenden Open-Source-Datensätzen. Er enthält zuvor ungenutzte Quellen wie PDFs und ArXiv-Papiere und dient der Forschung im multimodal …

README

--- license: cc-by-4.0 task_categories: - image-to-text - text-generation language: - en tags: - multimodal pretty_name: MINT-1T size_categories: - 100B<n<1T configs: - config_name: data-v1.1 data_files: - split: train path: data_v1_1/*.parquet --- <h1 align="center"> 🍃 MINT-1T:<br>Skalierung von …

查看完整页面 · 查看原文