oasst1 OpenAssistant

OpenAssistant Conversations Dataset (OASST1) Dataset Summary In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations (OASST1), a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 fully annotated conversation trees. The corpus is a product of a worldwide crowd-sourcing effort… See the full description on the dataset page: https://huggingface.co/datasets/OpenAssistant/oasst1.

Type
dataset
License
apache-2.0
Language
en
Downloads
21,688
Likes
1,559
Access
public
Files
0

Tags

  • 对话数据
  • 多语言语料
  • 指令微调
  • 大语言模型
  • 文本数据集
  • 人类反馈
  • 多语言

README

--- license: apache-2.0 dataset_info: features: - name: message_id dtype: string - name: parent_id dtype: string - name: user_id dtype: string - name: created_date dtype: string - name: text dtype: string - name: role dtype: string - name: lang dtype: string - name: review_count dtype: int32 - name…

查看完整页面 · 查看原文