Anthropic / hh-rlhf
人类偏好数据,用于训练帮助性和无害性奖励模型
仓库许可不等于数据内容权利已全部清理,商用仍需核对数据来源、隐私与版权条款。
项目资料与外部反馈
该数据集标注质量存在多项已知问题,但 MIT 许可明确允许商用,适合作为 RLHF 偏好建模的基线参考,不适合直接用于生产级对齐训练。
- RLHF 偏好建模与奖励模型训练
- DPO 对齐微调
- 对话助手对齐研究基线
- 数据集中存在 chosen 和 rejected 完全相同的样本对,标注一致性存在问题。
- 偏好数据的标注存在固有歧义和质量低下问题,会影响奖励模型训练效果。
- harmlessness 分类标准存在类别错误。
- 使用默认 load_dataset() 加载时需要指定子目录(data_dir),否则可能因 schema 不一致出错。
- 数据集预览功能无法正常工作,数据生成时可能抛出 DatasetGenerationCastError。
- 部分链接返回 404 错误。
来源
Anthropic/hh-rlhf · DiscussionsAnthropic/hh-rlhf · Does this HF dataset include all 5 folders from the github repoAnthropic/hh-rlhf · dataset preview does not workImproving Reinforcement Learning from Human Feedback Using Contrastive RewardsAnthropic HH-RLHF — Preference (RLHF / DPO) Dataset for LLM Training | LLM Configurator
截至 2026-07-22