monet jasperai

Dataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.

Type
dataset
License
apache-2.0
Language
en
Downloads
184,360
Likes
149
Access
public
Files
0

Tags

  • 多模态数据集
  • 文生图
  • 图像数据集
  • 图像特征提取
  • 合成数据
  • 数据处理
  • 计算机视觉
  • 零样本学习

README

--- license: apache-2.0 pretty_name: MONET task_categories: - text-to-image - image-feature-extraction - zero-shot-image-classification language: - en size_categories: - 100M<n<1B tags: - text-to-image - image-text - multimodal - captioning - synthetic-data configs: - config_name: parquet data_file…

查看完整页面 · 查看原文