stack-v3-train HuggingFaceCode

🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.

类型
dataset
许可
odc-by
语言
code
下载量
237,737
点赞
344
访问
public
文件
500

标签

  • 代码数据
  • 预训练语料
  • 文本生成
  • 代码生成
  • 大语言模型
  • NLP
  • 多语言语料

摘要

The Stack v3 是当前最大的开源源代码数据集,从 GitHub 抓取 2025 年 8 月状态的超过 173M 仓库、713 种语言、约 4.9 万亿 token 的源代码,训练子集共 15.9TB。它按整个仓库分组并以内联方式保存文件内容,专门用于带完整仓库上下文的代码大语言模型预训练,且支持 streaming 与 DuckDB 流式分析。该数据集用于训练代码模型,提升代码理解与生成能力,适用于预训练代码 LLM、按语言筛选及自建训练混合等场景。

README

--- thumbnail: "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png" annotations_creators: [] language_creators: - crowdsourced - expert-generated language: - code license: - odc-by multilinguality: - multilingual size_categories: - 100M<n<1B source_dataset…

查看完整页面 · 查看原文