stack-v3-train HuggingFaceCode
🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.
- Typ
- dataset
- Lizenz
- odc-by
- Sprache
- code
- Downloads
- 237,737
- Likes
- 344
- Zugriff
- public
- Dateien
- 500
Tags
- 代码数据集
- 预训练语料
- 多语言语料
- 文本生成
- 大语言模型
- 机器学习
- Python
Zusammenfassung
Der Stack v3 ist der größte und aktuellste offene Quellcode-Datensatz (Stand August 2025), der direkt von GitHub gecrawlt und speziell zum Pre-Training von Code-LLMs mit vollem Repository-Kontext erstellt wurde. Der Trainings-Subset umfasst 15,9 TB Quellcode aus 173M Repositories in 713 Programmier…
README
--- thumbnail: "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png" annotations_creators: [] language_creators: - crowdsourced - expert-generated language: - code license: - odc-by multilinguality: - multilingual size_categories: - 100M<n<1B source_dataset…