AI & Technology — 2026-03-26

Google TurboQuant Compresses AI Memory by 6x, Rattles Memory Chip Stocks

Google unveiled TurboQuant, a compression method that reduces LLM key-value cache memory by at least 6x—quantizing from 16 bits to just 3 bits per value with no accuracy loss and no retraining required. The technique combines PolarQuant (vector-to-polar coordinate transformation, to be presented at AISTATS 2026) and Quantized Johnson-Lindenstrauss (QJL) error correction. On NVIDIA H100 accelerators, the 4-bit implementation achieved up to 8x attention speedup versus 32-bit unquantized baselines. Tested on Gemma, Mistral, and Llama-3.1-8B models across five long-context benchmarks. The paper will be formally presented at ICLR 2026 in April. SK Hynix fell 6.2% and Samsung dropped 4.8% on the Korea Exchange; U.S.-listed Micron and SanDisk also declined. Morgan Stanley called the sell-off excessive, arguing the technique only affects inference workloads—not training, which drives the bulk of HBM procurement—and remains a lab result without production deployment.

Analysis
The 6x memory reduction addresses inference cost, the largest operational expense for LLM deployments, but does nothing for training, which drives the bulk of GPU and HBM procurement. The chip stock sell-off may be overdone: training workloads still require maximum memory bandwidth, and TurboQuant remains a lab result without production deployment.
3 sources
  1. A Google AI breakthrough is pressuring memory chip stocks from Samsung to Micron - CNBC
  2. Google's new TurboQuant algorithm speeds up AI memory 8x, cutting costs by 50% or more - VentureBeat
  3. TurboQuant: Redefining AI efficiency with extreme compression - Google Research Blog

View in full brief →

UNCLASSIFIED // OPEN SOURCE