Google TurboQuant Compresses AI Memory by 6x, Rattles Memory Chip Stocks
Google unveiled TurboQuant, a compression method that reduces LLM key-value cache memory by at least 6x—quantizing from 16 bits to just 3 bits per value with no accuracy loss and no retraining required. The technique combines PolarQuant (vector-to-polar coordinate transformation, to be presented at AISTATS 2026) and Quantized Johnson-Lindenstrauss (QJL) error correction. On NVIDIA H100 accelerators, the 4-bit implementation achieved up to 8x attention speedup versus 32-bit unquantized baselines. Tested on Gemma, Mistral, and Llama-3.1-8B models across five long-context benchmarks. The paper will be formally presented at ICLR 2026 in April. SK Hynix fell 6.2% and Samsung dropped 4.8% on the Korea Exchange; U.S.-listed Micron and SanDisk also declined. Morgan Stanley called the sell-off excessive, arguing the technique only affects inference workloads—not training, which drives the bulk of HBM procurement—and remains a lab result without production deployment.
The 6x memory reduction addresses inference cost, the largest operational expense for LLM deployments, but does nothing for training, which drives the bulk of GPU and HBM procurement. The chip stock sell-off may be overdone: training workloads still require maximum memory bandwidth, and TurboQuant remains a lab result without production deployment.