fleurs google

FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.

유형
dataset
라이선스
cc-by-4.0
언어
afr
다운로드
97,349
좋아요
443
접근
public
파일
0

태그

  • 음성 데이터셋
  • 음성 인식
  • 다국어
  • 평가 데이터셋
  • 음성 데이터
  • 벤치마크
  • 다중언어

요약

FLEURS는 FLoRes 기계번역 벤치마크의 음성 버전으로, 102개 언어에 걸친 병렬 음성-텍스트 데이터셋입니다. 음성 인식, 번역, 분류, 검색 등 다양한 음성 태스크를 평가하는 XTREME-S 벤치마크의 기반이며, 2009개 문장과 약 10시간의 학습 데이터를 포함합니다. 다국어 음성 인식 모델을 학습·평가하는 데 널리 사용되며, 여러 어족과 언어 그룹을 포괄합니다.

README

--- annotations_creators: - expert-generated - crowdsourced - machine-generated language_creators: - crowdsourced - expert-generated language: - afr - amh - ara - asm - ast - azj - bel - ben - bos - cat - ceb - cmn - ces - cym - dan - deu - ell - eng - spa - est - fas - ful - fin - tgl - fra - gle …

查看完整页面 · 查看原文