finephrase HuggingFaceFW

Dataset Card for HuggingFaceFW/finephrase Dataset Summary Synthetic data generated by DataTrove: Model: HuggingFaceTB/SmolLM2-1.7B-Instruct (main) Source dataset: HuggingFaceFW/fineweb-edu, config sample-350BT, split train Generation config: temperature=1.0, top_p=1.0, top_k=50, max_tokens=2048, model_max_context=8192 Speculative decoding: {"method":"suffix","num_speculative_tokens":32} System prompt: None Input column: text Prompt families: faq prompt Rewrite… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finephrase.

種別
dataset
ライセンス
odc-by
言語
en
ダウンロード
252,026
いいね
143
アクセス
public
ファイル
0

タグ

  • 合成データ
  • 事前学習コーパス
  • 大規模データセット
  • 大言語モデル
  • テキスト生成
  • 英語テキストデータ
  • 命令微調整
  • DataTrove

概要

SmolLM2-1.7B-InstructでFineWeb-Eduの約3.4億文書をFAQ・数学・表・チュートリアルの4形式に書き換えて生成した大規模合成データセットです。約13.5億サンプル・486Bトークンを収録し、言語モデルの事前学習や指令微調整向けの高品質データ供給を目的とします。DataTroveパイプラインで生成され、odc-byライセンスで公開されています。

README

--- language: - en license: odc-by tags: - SmolLM2-1.7B-Instruct - fineweb-edu - synthetic - datatrove annotations_creators: - machine-generated language_creators: - found pretty_name: HuggingFaceFW/finephrase size_categories: - n>1M source_datasets: - HuggingFaceFW/fineweb-edu/sample-350BT task_ca…

查看完整页面 · 查看原文