chatbot_arena_conversations lmsys
Chatbot Arena Conversations Dataset This dataset contains 33K cleaned conversations with pairwise human preferences. It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023. Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp. To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.
- 类型
- dataset
- 许可
- cc
- 下载量
- 2,916
- 点赞
- 477
- 访问
- gated
- 文件
- 0
标签
- 对话数据
- 评测基准
- 人类反馈
- 大语言模型
- 多语言语料
- 文本数据集
摘要
该数据集收录了2023年4-6月间来自13000个独立IP的33K条清洗后聊天机器人对战对话,每条样本包含两个模型的完整对话、用户投票、去匿名化用户ID及语言/毒性检测标签。它服务于大语言模型的人类偏好对齐评估,是构建奖励模型、评测模型回答质量及进行RLHF训练的重要基准数据,适合研究对话质量判断与模型排名场景。