OpenThoughts-1k-sample ryanmarten
[!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.
- Typ
- dataset
- Downloads
- 1,589,808
- Likes
- 49
- Zugriff
- public
- Dateien
- 0
Tags
- Synthetische Daten
- Pre-training-Abdeckung
- Code-Daten
- Instruction-Tuning
- Große Sprachmodelle
- Reasoning
- Textdatensatz
- Feinabstimmung
Zusammenfassung
Dieses Dataset ist eine 1.000-Beispiele-Stichprobe des OpenThoughts-114k-Datensatzes. Es enthält hochwertige synthetische Reasoning-Daten aus DeepSeek-R1, die Mathematik, Wissenschaft, Code und Rätsel abdecken und zur Feinabstimmung der OpenThinker-Modelle verwendet werden. Geeignet für das Trainin…
README
--- configs: - config_name: default data_files: - split: train path: data/train-* - config_name: metadata data_files: - split: train path: metadata/train-* dataset_info: - config_name: default features: - name: system dtype: string - name: conversations list: - name: from dtype: string - name: valu…