the-stack-v2 bigcode
The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset <-- you are here bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2.
- 类型
- dataset
- 许可
- other
- 语言
- code
- 下载量
- 12,690
- 点赞
- 615
- 访问
- gated
- 文件
- 0
标签
- 代码数据
- 预训练语料
- 代码生成
- 大语言模型
- 多语言语料
- 文本数据集
摘要
The Stack v2 是 BigCode 推出的大型多语言代码预训练语料数据集,覆盖600多种编程语言,并提供多个去重与过滤版本(如 dedup、train-full-ids、train-smol-ids),满足不同规模的训练需求。该数据集主要服务于大语言模型的代码生成能力训练,是代码大模型与代码智能领域的重要基础资源。