Prefer to read without ads? Become a member — from $10/month — and support the work. Already a member? Log in to read ad-free on this device.

SRAM Scaling: Beginning of the End?

The impending crisis in hardware design that no one is talking about

Moore’s Law, like Amdahl’s Law, always has been more of an observation than a law. It stemmed from Gordon Moore, co-founder of Intel, observing in 1965 that with process improvements, transistor density was doubling every year. Later, around 1975, he revised that estimate downward to a doubling every 2 years—still a prodigious growth rate. And ever since, there has been a drumbeat of predictions that Moore’s Law would come to an end, and ruminations on what would come next after that happened.

Some perceived barriers, like the 1-micron barrier (or in today’s parlance, the 1,000-nm barrier) turned out essentially to be psychological, like the 4-minute mile. The perception that Moore’s Law might come to an unwelcome end has motivated investments in alternative technologies, like Josephson junctions or gallium arsenide as a substrate, for decades.

For the last few years, as process improvements continued to roll out (albeit more unevenly than they did in the halcyon 1990s), an alarming trend started to be observed: logic was benefiting markedly more from improved scaling than SRAM. In December 2022, Tom’s Hardware published “TSMC’s 3nm Node: No SRAM Scaling Implies More Expensive CPUs and GPUs,” and in February 2024, semiengineering.com published: “SRAM Scaling Issues, And What Comes Next.”

Why should we care? What is SRAM to you and me?

Well, for more than 30 years, the majority of the transistor budget allocated to CPUs has been SRAM in the form of caches. Modern CPUs have dozens of cores and each core has a very fast L1 (usually 64K each for code and data) and a moderately fast L2; the L3, or “LLC” (last level cache), or “uncore,” consists of an L3 cache that arbitrates external memory traffic, be it across the cache coherency interconnect or the integrated memory controllers on the CPU die. In short, the caches are comprised of SRAM and are designed to make memory seem faster (lower latency). SRAM has been such a dominant percentage of CPUs’ transistor budgets for the last 30 years that even significant instruction set improvements such as MMX (52 new instructions, c. 1996) and x64 (c. 2002) cost less than 10% total die area.

We’ve had scares like this before. A little over 20 years ago, a big inflection point in the history of Moore’s Law occurred: the end of Dennard scaling, where improvements in transistor density also delivered improvements in clock rates. The 1990s had been a decade of free beer, with Intel and AMD leading the charge from 25 MHz 80486 chips to 1,000-MHz chips by the end of the decade. In 2002, then-Intel VP Pat Gelsinger observed that it was becoming infeasible to parlay improvements in density into higher clock rates, and predicted multicore CPUs as a likely answer to the questions posed by this technical challenge. Intel introduced its first multicore in the Pentium Duo (c. 2006), and today, modern CPUs have many dozens of cores and peak performance is impossible to attain without multithreading.

But with benefit of hindsight, we now see that the primary beneficiary of the end of Dennard scaling was not any CPU vendor, but NVIDIA, with its GPUs and the CUDA technology underpinning their AI hardware business. NVIDIA GPUs, the story went, were designed for throughput, not latency; so they were able to target lower clock rates and, instead of caches designed to reduce the effective latency of memory, the SRAM in GPUs is allocated to:

  • The humongous register file (largest memory on the chip), SRAM that helps cover memory latencies and instruction latencies via thread level parallelism (TLP);

  • Shared memory, a software-managed cache, occupies a significant portion of each Streaming Multiprocessor (SM), of which there are 140 on a modern GPU;

  • The L1 caches for each SM, which typically are aliased on to the shared memory (and the proportion of L1/shared memory is under software control); and

  • the L2 cache, dozens of megabytes, is yet more SRAM.

The design of each of these hardware units reflects the different types of memory traffic serviced by the SRAM.

But the point is, the transistor budgets of both CPUs and GPUs is dominated by SRAM.

And for a scary few years, SRAM did not seem to benefiting from scaling to the same degree as logic. Quoting from the above-linked semiengineering.com article:

“In traditional scaling of planar devices, gate length and gate oxide thickness were scaled down together to improve performance and control of the short-channel effect. A thinner oxide enabled the performance gain in lower VDD level, which is advantageous for SRAM in reducing both leakage and dynamic power,” said Jongsin Yun, memory technologist at Siemens EDA. “However, in recent technology node migrations, we’ve barely seen further scaling oxide or VDD levels. Moreover, the geometric shrinkage of transistors results in thinner metal interconnects, leading to increased parasitic resistance and, consequently, more power loss, and RC delay. As AI design increasingly demands more internal memory access, it has become a significant challenge for SRAM to further scale its power and performance benefits in technology node migration.”

What does this mean?

It’s as big a question for hardware and software designers as the end of Dennard scaling, or the end of Moore’s Law itself. One of the more obvious consequences is that, coupled with advances in packaging technology, chiplets that are especially RAM-intensive (e.g., I/O dies with cache) can target less-dense, more-economical fabrication processes. As the article mentions, AMD’s 3D V-Cache technology “allows additional SRAM cache memory to be stacked on top of processors, increasing the amount of cache available to the processor cores.”

A less-obvious innovation would be to trade compute for effective bandwidth; can tightly-integrated compression technology reduce the SRAM footprint needed to service a given memory traffic load?

Of course, the staggering investments in continued advancement of semiconductor fabrication technology, which have fueled that ongoing innovation, may forestall the apocalyptic scenario.

TSMC is hiring more memory designers to improve SRAM density, but whether they can eke more out of SRAM remains to be seen. “Sometimes you can make things better by applying more people, but only up to a point,” said Tate. “Over time, customers will need to think about architectures that don’t use SRAM as intensively as they do now.”

And tentatively, TSMC’s investments may have paid off: the February 15, 2024 article is followed by the November 2024 “SRAM Scaling Isn’t Dead After All.” All it will cost you is the very latest and most expensive TSMC process!

Moore’s Law is bound to run out sometime. To be honest, it has outlasted everyone’s reasonable expectation! We keep getting previews of what it might be like, and having to design hardware (and, sometimes, the software that runs on it) accordingly. Without impending physical laws to constrain and inform our designs, what would we do to keep our industry fun and exciting?