stack-v3-train HuggingFaceCode

🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.

Tipo
dataset
Licencia
odc-by
Lenguaje
code
Descargas
237,737
Me gusta
344
Acceso
public
Archivos
500

Etiquetas

  • código fuente
  • preentrenamiento
  • texto
  • preentrenamiento de LLM
  • dataset de código
  • GitHub
  • NLP
  • multilingüe

Resumen

The Stack v3 es el dataset abierto de código fuente más grande y actualizado, extraído de GitHub (estado de agosto 2025), con 15,9 TB y ~4,9 billones de tokens en 713 lenguajes. Está diseñado para preentrenar LLM de código con contexto de repositorio completo, agrupando los archivos y su contenido …

README

--- thumbnail: "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png" annotations_creators: [] language_creators: - crowdsourced - expert-generated language: - code license: - odc-by multilinguality: - multilingual size_categories: - 100M<n<1B source_dataset…

查看完整页面 · 查看原文