Hugging Face Blog·· 2023-02-15AI 评分26
我们为何转向 Hugging Face Inference Endpoints,或许你也应该考虑
Why we’re switching to Hugging Face Inference Endpoints, and maybe you should too
AI 导读
Hugging Face 推出托管服务 Inference Endpoints,可将 Hub 上的模型部署到 AWS、Azure(GCP 即将支持)等多种实例类型上。作者将原本跑在 AWS ECS + Fargate 上的 CPU 推理模型迁移过去,用 RoBERTa 文本分类模型测试显示,large 实例延迟约 80ms,比原方案约 200ms 快一倍以上,但成本高出 24% 至 50%。
来源:Hugging Face Blog · huggingface.co