OpenThoughts-1k-sample ryanmarten

[!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.

Typ
dataset
Downloads
1,589,808
Likes
49
Zugriff
public
Dateien
0

Tags

  • Synthetische Daten
  • Pre-training-Abdeckung
  • Code-Daten
  • Instruction-Tuning
  • Große Sprachmodelle
  • Reasoning
  • Textdatensatz
  • Feinabstimmung

Zusammenfassung

Dieses Dataset ist eine 1.000-Beispiele-Stichprobe des OpenThoughts-114k-Datensatzes. Es enthält hochwertige synthetische Reasoning-Daten aus DeepSeek-R1, die Mathematik, Wissenschaft, Code und Rätsel abdecken und zur Feinabstimmung der OpenThinker-Modelle verwendet werden. Geeignet für das Trainin…

README

--- configs: - config_name: default data_files: - split: train path: data/train-* - config_name: metadata data_files: - split: train path: metadata/train-* dataset_info: - config_name: default features: - name: system dtype: string - name: conversations list: - name: from dtype: string - name: valu…

查看完整页面 · 查看原文