OpenOrca Open-Orca
🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.
- 类型
- dataset
- 许可
- mit
- 语言
- en
- 下载量
- 20,179
- 点赞
- 1,582
- 访问
- public
- 文件
- 0
标签
- 指令微调
- 对话数据
- 文本数据集
- 大语言模型
- NLP
- 数据处理
- 推理
摘要
OpenOrca是基于FLAN Collection扩充的增强指令数据集,包含约1M条GPT-4和3.2M条GPT-3.5的完成结果,其分布与Orca论文对齐。该数据核心价值在于用大规模语言模型生成逐步推理轨迹来增强数据,助力训练出在推理任务上表现出色的小规模开源模型(如Mistral-7B-OpenOrca)。主要面向需要高质量指令微调与推理能力语料的NLP研究人员和开发者。
README
--- language: - en license: mit task_categories: - conversational - text-classification - token-classification - table-question-answering - question-answering - zero-shot-classification - summarization - feature-extraction - text-generation - text2text-generation pretty_name: OpenOrca size_categori…