跳到正文
原文
Hugging Face Blog·· 2023-05-16AI 评分28

BigCode 背后的大规模近重复数据删除

Large-scale Near-deduplication Behind BigCode

AI 导读

Hugging Face 博客分享了 BigCode 项目在大规模代码数据集上做近重复数据删除(near-deduplication)的实践,采用 MinHash + LSH 方法,参数为 (256, 0.7, 5)。该工作延续自 BigScience 的 ROOTS 语料去重经验,此前已确认去重能提升代码模型表现,同时使用更小的数据集。

来源:Hugging Face Blog · huggingface.co