databricks-dolly-15k databricks
Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several of the behavioral categories outlined in the InstructGPT paper, including brainstorming, classification, closed QA, generation, information extraction, open QA, and summarization. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Creative Commons Attribution-ShareAlike 3.0 Unported… See the full description on the dataset page: https://huggingface.co/datasets/databricks/databricks-dolly-15k.
- 유형
- dataset
- 라이선스
- cc-by-sa-3.0
- 언어
- en
- 다운로드
- 43,318
- 좋아요
- 1,075
- 접근
- public
- 파일
- 0
태그
- 명령어 미세조정
- 지시 데이터셋
- 대화 데이터
- 합성 데이터
- LLM
- 질문응답
- NLP
요약
이 데이터셋은 수천 명의 Databricks 직원이 생성한 15,000개 이상의 지시-응답 쌍으로 구성된 오픈소스 명령어 데이터셋입니다. InstructGPT 논문의 행동 범주(브레인스토밍, 분류, 폐쇄형/개방형 QA, 생성, 정보 추출, 요약)를 다룹니다. 대형 언어 모델의 명령어 미세조정 및 합성 데이터 생성, 데이터 증강에 활용할 수 있으며 CC BY-SA 3.0 라이선스로 상업적 사용도 가능합니다.
README
--- license: cc-by-sa-3.0 task_categories: - question-answering - summarization language: - en size_categories: - 10K<n<100K --- # 요약 `databricks-dolly-15k`는 수천 명의 Databricks 직원들이 [InstructGPT](https://arxiv.org/abs/2203.02155) 논문에 설명된 여러 행동 범주(브레인스토밍, 분류, 폐쇄형 QA, 생성, 정보 추출, 개방형 QA, 요약)에서 생성한 지시 수행…