stack-v3-train HuggingFaceCode
🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.
- 유형
- dataset
- 라이선스
- odc-by
- 언어
- code
- 다운로드
- 237,737
- 좋아요
- 344
- 접근
- public
- 파일
- 500
태그
- 코드 데이터셋
- 사전학습 말뭉치
- 코드 생성
- 대규모 언어 모델
- 다국어
- 합성 데이터
- 데이터셋
요약
The Stack v3는 GitHub에서 크롤링한 최대 규모의 최신 오픈소스 소스코드 데이터셋으로, 전체 저장소 맥락에서 코드 LLM을 사전학습하기 위해 구축되었습니다. 파일 내용이 인라인으로 포함되어 다운로드 직후 바로 학습이 가능하며, 약 15.9TB의 소스코드와 713개 프로그래밍 언어, 1.73억 개 저장소를 담고 있습니다. 코드 모델 사전학습, 전체 저장소 맥락 학습, 코드 이해·생성 모델 연구에 적합합니다.
README
--- thumbnail: "https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train/resolve/main/assets/banner.png" annotations_creators: [] language_creators: - crowdsourced - expert-generated language: - code license: - odc-by multilinguality: - multilingual size_categories: - 100M<n<1B source_dataset…