UltraChat openbmb

Dataset Card for Dataset Name Dataset Description An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts. To ensure generation quality, two separate ChatGPT Turbo APIs are adopted in generation, where one plays the role of the user to generate queries and the other generates the response. We instruct the user model with… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraChat.

種別
dataset
ライセンス
mit
言語
en
ダウンロード
4,671
いいね
500
アクセス
public
ファイル
0

タグ

  • 対話データ
  • 大規模言語モデル
  • テキスト生成
  • 指令微調整
  • 合成データ
  • 英語コーパス
  • NLP

概要

オープンソースかつ大規模なマルチターン対話データセットで、ChatGPT Turbo APIを利用して高品質な会話データを自動生成しています。ユーザー役と応答役に2つの別々のAPIを使い、人間らしい振る舞いを再現。世界知識、文章生成・創作、既存資料の支援という3セクターで構成され、対話型言語モデルの訓練に適しています。MITライセンスで英語データ約数百万件規模です。

README

--- license: mit task_categories: - conversational - text-generation language: - en size_categories: - 1M<n<10M pretty_name: UltraChat --- # Dataset Card for Dataset Name ## データセットの説明 オープンソースで大規模な、Turbo API を活用したマルチターン対話データです。プライバシー保護などの要素を考慮し、**当社はインターネット上にあるデータをプロンプトとして直接使用することはありません**。 生成品質を確保する…

查看完整页面 · 查看原文