PersonaHub proj-persona
Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/proj-persona/PersonaHub.
- 类型
- dataset
- 许可
- cc-by-nc-sa-4.0
- 语言
- en
- 下载量
- 9,861
- 点赞
- 792
- 访问
- public
- 文件
- 0
标签
- 合成数据
- 大语言模型
- 指令微调
- 数学推理
- 文本生成
- 生成式AI
- 数据处理
摘要
PersonaHub 提出一种"人设驱动"的数据合成方法,利用大语言模型中的多样视角来生成海量多样化合成数据,包含自动从网络中精选的10亿种人设(约占世界人口的13%)。该数据集提供数学与逻辑推理问题、指令、知识文本、游戏NPC及工具函数等多种合成数据样本,并发布20万人设预览和3.7亿精英人设,适用于规模化合成高质量训练数据以推动LLM研究与发展。
README
--- license: cc-by-nc-sa-4.0 task_categories: - text-generation - text-classification - token-classification - fill-mask - table-question-answering - text2text-generation language: - en - zh tags: - synthetic - text - math - reasoning - instruction - tool - persona size_categories: - 100M<n<1B conf…