NatureBench JoeLiu996

Dataset Card for NatureBench NatureBench is a cross-discipline benchmark of 27 tasks distilled from peer-reviewed Nature-family publications, spanning 6 scientific domains. It is designed to evaluate whether AI coding agents can move beyond reproduction toward discovery: each task asks an agent to solve a real scientific machine-learning problem and is scored against the source paper's reported state of the art. 📄 arXiv paper: https://arxiv.org/abs/2606.24530 💻 GitHub code… See the full description on the dataset page: https://huggingface.co/datasets/JoeLiu996/NatureBench.

種別
dataset
ライセンス
other
言語
en
ダウンロード
255,741
いいね
0
アクセス
public
ファイル
0

タグ

  • ベンチマーク
  • 評価データ
  • エージェント
  • 科学機械学習
  • コード
  • コード生成
  • AI安全性

概要

NatureBenchは、Nature系列の査読付き論文から抽出した27タスク・6科学領域をカバーするクロスディシプリンなベンチマークデータセットです。AIコーディングエージェントが論文の再現を超えて新規発見に到達できるかを評価することを目的とし、各タスクはソース論文の報告されたSOTAに対する相対ギャップでスコアリングされます。コンテナ化された隔離環境でWeb検索を無効化して評価し、ショートカット解を防ぐ検証機構も備えています。科学的機械学習問題の解決能力を測るコーディングエージェント評価用途に適しています。

README

--- language: - en license: other license_name: mit-with-third-party-data license_link: LICENSE pretty_name: NatureBench size_categories: - n<1K tags: - coding-agents - benchmark - scientific-machine-learning - nature configs: - config_name: default data_files: - split: train path: manifest.jsonl -…

查看完整页面 · 查看原文