OpenCodeReasoning nvidia

OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and technical details behind OpenCodeReasoning. Github Repo - Access the complete pipeline used to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/OpenCodeReasoning.

種別
dataset
ライセンス
cc-by-4.0
ダウンロード
14,854
いいね
550
アクセス
public
ファイル
0

タグ

  • 合成データ
  • コードデータ
  • コード生成
  • 命令微調整
  • 推論
  • 蒸留
  • データセット

概要

OpenCodeReasoningは、競技プログラミングの推論能力を蒸留するために構築された、最大規模の推論ベース合成コードデータセットです。735,255件のPythonサンプルを28,319問の競プロ問題から収録しています。コード生成モデルの教師あり微調整(SFT)に最適で、R1が生成した解を含み、CodeForces等複数のプラットフォーム由来の問題で構成されています。

README

--- license: cc-by-4.0 size_categories: - 100K<n<1M pretty_name: OpenCodeReasoning dataset_info: - config_name: split_0 features: - name: id dtype: string - name: input dtype: string - name: output dtype: string - name: source dtype: string - name: license dtype: string - name: dataset dtype: strin…

查看完整页面 · 查看原文