starcoderdata bigcode

StarCoder Training Dataset Dataset description This is the dataset used for training StarCoder and StarCoderBase. It contains 783GB of code in 86 programming languages, and includes 54GB GitHub Issues + 13GB Jupyter notebooks in scripts and text-code pairs, and 32GB of GitHub commits, which is approximately 250 Billion tokens. Dataset creation The creation and filtering of The Stack is explained in the original dataset, we additionally decontaminate and… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoderdata.

类型
dataset
许可
other
语言
code
下载量
29,185
点赞
533
访问
gated
文件
0

标签

  • 代码数据
  • 预训练语料
  • 代码生成
  • 大语言模型
  • 多语言
  • 文本数据集

摘要

该数据集是训练 StarCoder 和 StarCoderBase 的代码语料,包含 783GB 覆盖 86 种编程语言的代码,以及 GitHub Issues、Jupyter 笔记本与提交记录,约 2500 亿 token。它用于大规模预训练代码大语言模型,支持代码生成与补全任务,适合需要高质量多语言代码数据的模型训练场景。

查看完整页面 · 查看原文