The relentless pursuit of faster, more efficient AI computing has led engineers to explore radical designs, including stacking high-bandwidth memory (HBM) directly atop GPUs. While this "ultimate 3D integration" promises to slash memory bottlenecks, early simulations revealed a major obstacle: catastrophic overheating. However, recent research from Imec suggests that clever engineering could cool things down enough to make it feasible.

The Heat is On: Initial Findings

Currently, advanced GPUs from companies like AMD and Nvidia use a 2.5D packaging approach, where the GPU is placed alongside HBM chips on an interposer. This minimizes the distance between the processor and memory, crucial for AI applications that demand rapid data transfer. Imec's simulations of this setup showed a GPU consuming 414 watts, reaching a peak temperature of just under 70°C – typical for such a processor. "While this approach is currently used, it does not scale well for the future—especially as it blocks two sides of the GPU, limiting future GPU-to-GPU connections inside the package,” Yukai Chen, a senior researcher at Imec, told engineers at the IEEE International Electron Device Meeting (IEDM).

The initial 3D stacking model, however, painted a grim picture. Simply placing HBM chips directly on top of the GPU caused temperatures to skyrocket to 140°C, far exceeding the typical 80°C limit for GPUs. This thermal runaway threatened to render the entire concept unviable, demanding innovative solutions.

Engineering a Cool Solution

Undeterred, the Imec team embarked on a series of system technology co-optimizations. Their first move involved removing a redundant silicon layer – the base die in the HBM stack. By integrating the memory control circuits directly into the GPU, they eliminated the need for a separate data multiplexer. According to Imec's James Myers, this shift also freed up space on the GPU by removing the demultiplexing circuits, though it only resulted in a minor temperature reduction of 4°C.

Next, the team looked at the memory-bound nature of large language models. They theorized that the increased bandwidth from 3D stacking would allow them to reduce the GPU clock speed without sacrificing performance. By slowing the GPU by 50%, they achieved a significant 20°C temperature drop. Myers noted that increasing the clock frequency to 70% led to a GPU that was only 1.7 °C warmer.

Further optimizations included merging the HBM stacks to eliminate heat-trapping regions, thinning the top die of the stack, and filling surrounding space with conductive silicon. Finally, they implemented cooling on both the top and bottom of the package, resulting in a final temperature of approximately 70°C. This dual-sided cooling approach proved crucial, dropping the temperature by a significant 17°C.

While the research presented at IEDM demonstrates the potential feasibility of HBM-on-GPU designs, Myers cautions against premature conclusions. “We are simulating other system configurations to help build confidence that this is or isn’t the best choice,” he says. “GPU-on-HBM is of interest to some in industry,” because it puts the GPU closer to the cooling. But it would likely be a more complex design, because the GPU’s power and data would have to flow vertically through the HBM to reach it.

"We are simulating other system configurations to help build confidence that this is or isn’t the best choice."

— James Myers, Imec

The implications of this research extend beyond just cooling techniques. It highlights the increasing importance of system-level thinking in chip design. As Moore's Law continues to slow, innovations in packaging and thermal management will be critical to unlocking further performance gains. The coming years will be crucial in determining whether 3D-stacked GPUs become a mainstream reality, and which approach will be the most practical for the industry to implement.