OpenOrca Open-Orca
🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.
- Typ
- dataset
- Lizenz
- mit
- Sprache
- en
- Downloads
- 20,179
- Likes
- 1,582
- Zugriff
- public
- Dateien
- 0
Tags
- Synthetische Daten
- Feinabstimmung
- Große Sprachmodelle
- Englisch
- Textgenerierung
- NLP
- Vortrainingskorpus
Zusammenfassung
OpenOrca ist ein großer augmentierter Datensatz auf Basis der FLAN-Collection, der etwa 1M GPT-4- und 3,2M GPT-3.5-Antworten mit Reasoning-Traces enthält. Er wird hauptsächlich für das Training und die Evaluierung leistungsfähiger LLM-Checkpoints in der natürlichen Sprachverarbeitung genutzt und ha…
README
--- language: - en license: mit task_categories: - conversational - text-classification - token-classification - table-question-answering - question-answering - zero-shot-classification - summarization - feature-extraction - text-generation - text2text-generation pretty_name: OpenOrca size_categori…