OpenOrca Open-Orca

🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.

Typ
dataset
Lizenz
mit
Sprache
en
Downloads
20,179
Likes
1,582
Zugriff
public
Dateien
0

Tags

  • Synthetische Daten
  • Feinabstimmung
  • Große Sprachmodelle
  • Englisch
  • Textgenerierung
  • NLP
  • Vortrainingskorpus

Zusammenfassung

OpenOrca ist ein großer augmentierter Datensatz auf Basis der FLAN-Collection, der etwa 1M GPT-4- und 3,2M GPT-3.5-Antworten mit Reasoning-Traces enthält. Er wird hauptsächlich für das Training und die Evaluierung leistungsfähiger LLM-Checkpoints in der natürlichen Sprachverarbeitung genutzt und ha…

README

--- language: - en license: mit task_categories: - conversational - text-classification - token-classification - table-question-answering - question-answering - zero-shot-classification - summarization - feature-extraction - text-generation - text2text-generation pretty_name: OpenOrca size_categori…

查看完整页面 · 查看原文