跳到正文
原文
Google Research·· 2026-03-25精选AI 评分62

Google 发布 TurboQuant 等三种向量量化算法,将 KV cache 压缩至 3 bit

TurboQuant: Redefining AI efficiency with extreme compression

AI 导读

Google 发布 TurboQuant 压缩算法,并配套提出 QJL 和 PolarQuant,用于解决向量量化中的内存开销问题,三篇论文将分别亮相 ICLR 2026 和 AISTATS 2026。

推荐理由

Google 提出 TurboQuant 等三种向量量化算法,可在不损失精度的前提下把 KV cache 压到 3 bit,读者可了解其原理与基准表现。

深读指南

适合谁读:关注大模型推理优化、KV cache 压缩与向量检索的工程与研究人员;需要了解量化算法原理与实验结论的技术读者。

  • 传统向量量化本身会带来额外内存开销,因为需要为每个数据块以全精度存储量化常数,可能每数多出 1 到 2 bit。

    This overhead can add 1 or 2 extra bits per number, partially defeating the purpose of vector quantization.
  • TurboQuant 采用两阶段思路:先用 PolarQuant 做主要压缩,再用仅 1 bit 的 QJL 处理残差误差以消除偏差。

    TurboQuant uses a small, residual amount of compression power (just 1 bit) to apply the QJL algorithm to the tiny amount of error left over from the first stage.
  • PolarQuant 通过将向量转为极坐标,利用角度分布已知且集中的特点,省去昂贵的数据归一化步骤,从而消除传统方法的内存开销。

    the model no longer needs to perform the expensive data normalization step because it maps data onto a fixed, predictable "circular" grid

局限与前提:本文为 Google 官方发布博客,实验细节、数据集配置与理论证明均未在正文展开,仅给出聚合图表与结论性描述;所引用的基准(LongBench、Needle In A Haystack、ZeroSCROLLS、RULER、L-Eval)与对比基线(KIVI、PQ、RabbiQ)的具体设置、硬件条件与统计显著性未说明,因此“零精度损失”“至少 6x 压缩”“最高 8x 加速”等应视为作者报告的结果而非独立验证结论。此外,正文未给出可复现步骤、超参数或失败案例,实际部署效果需以论文与后续复现为准。

对照原文阅读

来源:Google Research · research.google