finephrase HuggingFaceFW
Dataset Card for HuggingFaceFW/finephrase Dataset Summary Synthetic data generated by DataTrove: Model: HuggingFaceTB/SmolLM2-1.7B-Instruct (main) Source dataset: HuggingFaceFW/fineweb-edu, config sample-350BT, split train Generation config: temperature=1.0, top_p=1.0, top_k=50, max_tokens=2048, model_max_context=8192 Speculative decoding: {"method":"suffix","num_speculative_tokens":32} System prompt: None Input column: text Prompt families: faq prompt Rewrite… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finephrase.
- 種別
- dataset
- ライセンス
- odc-by
- 言語
- en
- ダウンロード
- 252,026
- いいね
- 143
- アクセス
- public
- ファイル
- 0
タグ
- 合成データ
- 事前学習コーパス
- 大規模データセット
- 大言語モデル
- テキスト生成
- 英語テキストデータ
- 命令微調整
- DataTrove
概要
SmolLM2-1.7B-InstructでFineWeb-Eduの約3.4億文書をFAQ・数学・表・チュートリアルの4形式に書き換えて生成した大規模合成データセットです。約13.5億サンプル・486Bトークンを収録し、言語モデルの事前学習や指令微調整向けの高品質データ供給を目的とします。DataTroveパイプラインで生成され、odc-byライセンスで公開されています。
README
--- language: - en license: odc-by tags: - SmolLM2-1.7B-Instruct - fineweb-edu - synthetic - datatrove annotations_creators: - machine-generated language_creators: - found pretty_name: HuggingFaceFW/finephrase size_categories: - n>1M source_datasets: - HuggingFaceFW/fineweb-edu/sample-350BT task_ca…