跳到正文
原文
Hugging Face Blog·· 2026-05-14AI 评分46

Hugging Face 详解连续批处理中的异步化:如何让 CPU 与 GPU 并行以提升 LLM 推理吞吐

Unlocking asynchronicity in continuous batching

AI 导读

Hugging Face 发文解析连续批处理中的异步化,指出默认同步批处理下 CPU 与 GPU 轮流工作,用 8B 模型、batch size 32 生成 8K tokens 时总耗时 300.6 秒,其中 24.0% 时间 GPU 处于空闲等待 CPU。

来源:Hugging Face Blog · huggingface.co