databricks-dolly-15k databricks
Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.
- Typ
- dataset
- Lizenz
- cc-by-sa-3.0
- Sprache
- en
- Downloads
- 43,318
- Likes
- 1,075
- Zugriff
- public
- Dateien
- 0
Tags
- 指令微调
- Textdatensatz
- Englisch
- 对话数据
- NLP
- Große Sprachmodelle
- Synthetische Daten
Zusammenfassung
Dieser Datensatz enthält über 15.000 von Databricks-Mitarbeitern erzeugte Paare aus Anweisung und Antwort in acht Kategorien wie Brainstorming, Klassifikation, Zusammenfassung und Frage-Antwort. Er dient der Instruktionsfeinabstimmung großer Sprachmodelle, synthetischer Datengenerierung und Datener…
README
--- license: cc-by-sa-3.0 task_categories: - question-answering - summarization language: - en size_categories: - 10K<n<100K --- # Zusammenfassung `databricks-dolly-15k` ist ein Open-Source-Datensatz mit instruktionsbasierten Einträgen, die von Tausenden von Databricks-Mitarbeitern in mehreren der …