fleurs google
FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.
- 種別
- dataset
- ライセンス
- cc-by-4.0
- 言語
- afr
- ダウンロード
- 97,349
- いいね
- 443
- アクセス
- public
- ファイル
- 0
タグ
- 音声データセット
- 音声認識
- 多言語データセット
- 評価データ
- ベンチマーク
- 機械翻訳
- 多言語
- 音声処理
概要
FLEURSは、FLoRes機械翻訳ベンチマークを音声に応用した多言語音声認識データセットです。102言語・10以上の語族を網羅し、音声認識・翻訳・分類・検索の4つのタスクファミリーを評価します。XTREME-Sベンチマークの一部として、多言語音声表現モデルや音声認識モデルの評価・ファインチューニングに広く利用されています。話者分離されたtrain/dev/testの分割により、クロスリンガル転移評価にも適しています。
README
--- annotations_creators: - expert-generated - crowdsourced - machine-generated language_creators: - crowdsourced - expert-generated language: - afr - amh - ara - asm - ast - azj - bel - ben - bos - cat - ceb - cmn - ces - cym - dan - deu - ell - eng - spa - est - fas - ful - fin - tgl - fra - gle …