Hugging Face Blog·· 2026-05-14AI 评分46
Hugging Face 详解连续批处理中的异步化:如何让 CPU 与 GPU 并行以提升 LLM 推理吞吐
Unlocking asynchronicity in continuous batching
AI 导读
Hugging Face 发文解析连续批处理中的异步化,指出默认同步批处理下 CPU 与 GPU 轮流工作,用 8B 模型、batch size 32 生成 8K tokens 时总耗时 300.6 秒,其中 24.0% 时间 GPU 处于空闲等待 CPU。
来源:Hugging Face Blog · huggingface.co