Research Webzine of the KAIST College of Engineering since 2014
Fall 2026 Vol. 27Researchers from KAIST, Google Research, DeepMind, and NYU unveiled TurboQuant, a new compression method that cuts AI memory use by up to six times while maintaining performance, opening the door to faster, more efficient AI services.

TurboQuant is a quantization algorithm that combines a random projection-based 1-stage quantization method and a QJL-based 2-stage residual quantization to improve memory efficiency while correcting accuracy.
As artificial intelligence models continue to grow in size and capability, one of the biggest technical challenges has been their heavy memory demand. Larger models require more expensive hardware and often run more slowly, limiting the wider deployment of high-performance AI services.
To address this problem, a joint research team from KAIST, Google Research, DeepMind, and New York University developed TurboQuant, a new algorithm designed to significantly reduce AI memory usage while preserving model performance. The work highlights a promising path toward making advanced AI systems faster, cheaper, and more practical across a wide range of devices and computing environments.
The key idea behind TurboQuant is to store the numerical data used in AI models more efficiently. Modern AI systems rely on large volumes of highly precise numbers, which consume substantial memory. TurboQuant compresses these values into a smaller representation while retaining the information that matters most for model accuracy. In simple terms, it works like compressing a high-resolution image while keeping the visible quality nearly unchanged.
According to the research team, TurboQuant can reduce the memory required by AI models by up to six times while maintaining nearly the same level of accuracy. This is particularly important because memory bottlenecks are one of the main factors that slow down AI inference and increase deployment costs in real-world applications.
TurboQuant achieves this through a two-stage compression process. First, the data is transformed so that it becomes more evenly distributed, making it easier to compress efficiently. Then, the small residual errors introduced in the first stage are compressed again in a second step, minimizing overall information loss. This two-stage strategy allows TurboQuant to achieve both high compression efficiency and strong accuracy, outperforming more conventional approaches.
The impact of this technology could be broad. It may enable larger AI models to run on the same hardware, reduce the cost of delivering AI services, and expand the use of advanced AI on personal devices such as smartphones and laptops, as well as in large-scale data centers. In the longer term, it may also influence the semiconductor industry by shifting attention from simply increasing memory capacity to improving memory efficiency.
This research is also significant because it reflects a strong international collaboration in which Korean researchers played a direct role in developing a core AI algorithm. The team plans to continue working on technologies that make large-scale AI models more efficient, with the expectation that TurboQuant will help lay the groundwork for next-generation AI systems and high-performance semiconductor technologies.
SweepLED: Finding Hidden Cameras with Decoupled Illumination Sweeps
Read moreWhat If Plants Could Play?
Read moreThe Era of “Molecular Refining” : Filtering Crude Oil Without Boiling
Read moreBeyond Hearing: Earphones That Sense the Body
Read moreBridging Neuroscience and Engineering: Brain-inspired Network for Robust State Estimation
Read more