Hugging Face Blog·· 2023-05-16AI 评分28
BigCode 背后的大规模近重复数据删除
Large-scale Near-deduplication Behind BigCode
AI 导读
Hugging Face 博客分享了 BigCode 项目在大规模代码数据集上做近重复数据删除(near-deduplication)的实践,采用 MinHash + LSH 方法,参数为 (256, 0.7, 5)。该工作延续自 BigScience 的 ROOTS 语料去重经验,此前已确认去重能提升代码模型表现,同时使用更小的数据集。
来源:Hugging Face Blog · huggingface.co