UltraFeedback openbmb

Introduction GitHub Repo UltraRM-13b UltraCM-13b UltraFeedback is a large-scale, fine-grained, diverse preference dataset, used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN). We then use these prompts to query multiple LLMs (see Table for model lists) and generate 4 different responses for each prompt, resulting in a total of 256k samples. To… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraFeedback.

유형
dataset
라이선스
mit
언어
en
다운로드
5,838
좋아요
432
접근
public
파일
0

태그

  • 대화 데이터
  • 보상 모델
  • 인간 피드백
  • 명령어 미세조정
  • 강화 학습
  • 평가 데이터셋
  • 대형 언어 모델
  • 데이터셋

요약

UltraFeedback은 보상 모델과 비평 모델을 훈련하기 위해 설계된 대규모 고품질 선호도 데이터셋입니다. UltraChat, ShareGPT, Evol-Instruct 등 다양한 소스에서 수집된 약 6.4만 개 프롬프트에 17개 서로 다른 LLM이 각각 4개의 응답을 생성하여 총 25.6만 샘플을 구축했습니다. 세분화된 주석 지침을 바탕으로 GPT-4가 지시 준수, 진실성, 정직성, 유용성 4가지 측면에서 수치·서면 피드백을 제공합니다. RLHF 연구자가 약 100만 개의 비교 쌍을 구성하여 보상 모델(예: UltraRM)과…

README

--- license: mit task_categories: - text-generation language: - en size_categories: - 100K<n<1M --- ## 소개 - [GitHub 저장소](https://github.com/thunlp/UltraFeedback) - [UltraRM-13b](https://huggingface.co/openbmb/UltraRM-13b) - [UltraCM-13b](https://huggingface.co/openbmb/UltraCM-13b) UltraFeedback은 강력…

查看完整页面 · 查看原文