orca-agentinstruct-1M-v1 microsoft

Dataset Card This dataset is a fully synthetic set of instruction pairs where both the prompts and the responses have been synthetically generated, using the AgentInstruct framework. AgentInstruct is an extensible agentic framework for synthetic data generation. This dataset contains ~1 million instruction pairs generated by the AgentInstruct, using only raw text content publicly avialble on the Web as seeds. The data covers different capabilities, such as text editing, creative… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-agentinstruct-1M-v1.

種別
dataset
ライセンス
cdla-permissive-2.0
言語
en
ダウンロード
1,584
いいね
465
アクセス
public
ファイル
0

タグ

  • 合成データ
  • 指示微調整
  • 対話データ
  • 大規模言語モデル
  • Agent
  • 命令微調整
  • テキストデータ
  • 質問応答

概要

Microsoftが開発したAgentInstructフレームワークで合成生成された、約100万件の指示ペア(プロンプトと応答)からなるデータセットです。テキスト編集、クリエイティブライティング、コーディング、読解、QA、RAG、テキスト分類など多様な能力をカバーし、任意の基本LLMの指示チューニング(命令微調整)に利用できます。Mistral-7bへの適用でAGIEval、MMLU、GSM8K、BBH、AlpacaEvalなど多くのベンチマークで大幅な性能向上が確認されています。

README

--- language: - en license: cdla-permissive-2.0 size_categories: - 1M<n<10M task_categories: - question-answering dataset_info: features: - name: messages dtype: string splits: - name: creative_content num_bytes: 288747542 num_examples: 50000 - name: text_modification num_bytes: 346421282 num_examp…

查看完整页面 · 查看原文