MMMLU openai

Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science. We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/openai/MMMLU.

Tipo
dataset
Licencia
mit
Lenguaje
ar
Descargas
11,123
Me gusta
523
Acceso
public
Archivos
0

Etiquetas

  • Benchmark
  • Evaluación de modelos
  • multilingüe
  • NLP
  • LLM
  • evaluación de modelos
  • razonamiento
  • arabe

Resumen

MMMLU es un benchmark multilingüe que traduce el conjunto de prueba de MMLU a 14 idiomas mediante traductores humanos profesionales. Evalúa el conocimiento general de los modelos de IA en 57 categorías, desde nivel elemental hasta temas profesionales avanzados como derecho, física, historia y cienc…

README

--- task_categories: - question-answering configs: - config_name: default data_files: - split: test path: test/*.csv - config_name: AR_XY data_files: - split: test path: test/mmlu_AR-XY.csv - config_name: BN_BD data_files: - split: test path: test/mmlu_BN-BD.csv - config_name: DE_DE data_files: - s…

查看完整页面 · 查看原文