Key Takeaways:
- DFT is more and more crucial for detecting defects in reminiscence cells, TSVs, microbumps, and high-speed interfaces, whereas energy integrity and thermal results add additional check challenges.
- Marginal or getting old interconnects are caught by a mix of BiST, embedded screens, boundary scan, at-speed testing, redundancy/restore, and in-system monitoring, a posh hierarchy of testing.
- HBM is a reliability bottleneck, requiring better visibility into particular person bonds and TSVs, and extra refined SI/PI and restore mechanisms for latent failures in advanced 2.5D/3D AI techniques.
The excessive value of area failures in knowledge facilities is driving huge adjustments in design-for-test, enabling chipmakers to establish impending failures and root out interconnect and die-to-die weaknesses. That is particularly vital for high-bandwidth reminiscence (HBM), which is appearing because the frontier structure for proving out 3D manufacturing and testing methods.
HBM is well-known for the super advantages it delivers to data-intensive duties similar to AI, high-resolution graphics processing, and different purposes that want large knowledge processing and high-speed knowledge switch. Its tightly coupled, vertically stacked DRAM creates ultrawide communication channels (1024, 2048 bits) that outperform all different reminiscence sorts. However testing HBM faces super challenges as a consequence of its advanced construction, which turns into significantly more durable with every machine era.
“Now we’re getting ready for HBM 5 expertise, which is able to enhance that stack top to 24, which suggests much more knowledge going forwards and backwards. Lots of ones and zeros in very shut proximity with no development in footprint/space,” mentioned Faisal Goriawalla, director of product administration at Synopsys. “In order the pitch between these indicators is squeezed, you could have electrical interference challenges such because the sufferer aggressor state of affairs, the place one internet toggling might trigger the opposite nets in its neighborhood to additionally flip, which was not supposed to occur. These tight pitches make the DFT elements of testing, similar to analysis, rather more difficult in multi-die expertise.”
That is particularly evident with microbumps and bonded interconnects, which regularly can’t be accessed straight. “By advantage of those interconnects being inside a multi-die bundle, the testing that may be carried out is restricted,” mentioned Noam Brousard, vp of options engineering at proteanTecs. “So we use minuscule screens on the tip of the HBM that measure every sign and decide how shut it’s to failure, offering an eye fixed diagram per lane with very excessive accuracy.”
Such preventive methods have gotten more and more engaging to knowledge facilities working giant language fashions, whose interruption results in important losses.
In HBM3, 12 or extra DRAM chiplets are stacked atop a silicon interposer, interconnected by finely spaced through-silicon vias (TSVs) and linked by microbumps. As a lot effort goes into testing the TSVs, microbumps, and die-die interfaces as for the reminiscence cells themselves. For that reason, DFT architectures have turn into crucial for producers of HBM modules, together with SK hynix, Samsung, and Micron.
The magnitude of this problem can’t be overstated. “HBM testing is usually a main bottleneck as a result of complexity of check program improvement,” mentioned Quoc Phan, expertise enablement supervisor of 3DIC DFT and yield at Siemens EDA. “Creating the specialised fault fashions and check algorithms mandatory for detecting defects in TSVs, microbumps, and inter-die interfaces is an intricate course of requiring deep experience. Integrating these superior exams seamlessly with the host processor’s DFT structure and making certain dependable communication between them provides substantial layers of complexity to each check program improvement and the following debugging phases.”
With the upcoming transition to hybrid bonding for superior HBM4 processes, rather more of the main focus can be on making certain the standard of every interconnect bond. “As we’re stacking these gadgets, the coplanarity, the warpage, the bonding processes, and all the pieces that we’re doing to make these particular person bonds from a C4 bump to a copper- to-copper bump or a die-to-die connection — every certainly one of these bonds is crucial, mentioned Jack Lewis, CTO of Modus Test. “The quantity of exact info we are able to get about these particular person bonds and the efficiency of the bonding processes will assist our clients ramp and hit yield entitlement shortly.”
DFT’s mission is to allow a constant method of testing chips post-manufacturing. However with an HBM stack, it additionally must detect defects within the reminiscence cells and periphery, TSVs, microbumps, die-die interfaces, base die circuitry, package-level interconnects, and high-speed I/O interfaces to the accelerator such because the UCI Categorical.
Although the technique by anybody chipmaker is proprietary, relationships amongst corporations are shifting. “DFT is developed as proprietary IP by every machine producer, so it’s troublesome to touch upon particular implementations,” mentioned Jin Yokoyama, senior director of reminiscence product advertising and marketing at Advantest. “Nevertheless, trying ahead, the boundary of tasks between SoC finish customers and reminiscence distributors could turn into more and more blurred and sophisticated, significantly with respect to how DFT is outlined and partitioned.”
A part of the rationale for this alteration is the adoption of customized HBM for particular purposes.
Reconfiguring for customized HBM
The transition from HBM3 to HBM4 consists of the choice of a customized logic base die that replaces the DRAM-built controller die of earlier generations. Customized HBM is engaging as a result of it permits the AI accelerator or GPU designer to optimize the reminiscence stack for its particular workload. That is significantly vital for AI coaching and inference, the place efficiency is extra typically bandwidth-limited slightly than compute-limited. The draw back is that the longer improvement and qualification cycle is more likely to restrict customized HBM to the highest-volume purposes.

Fig. 1: The HBM DRAM stack. Supply: Synopsys
Importantly, customized HBM’s modifications alter the testing panorama. “Customized HBM offers SOC designers super flexibility to configure the logic base die the way in which they need,” mentioned Goriawalla. So it signifies that in case you are in a knowledge middle AI coaching atmosphere the place latency and throughput are crucial elements, you possibly can configure your HBM controller for these objectives. However in case you are in an AI inference kind of utility, the place space and energy are larger considerations, then you possibly can configure the HBM controller and logic otherwise. So this implies we’ve to be considerate about DFT based mostly on the use case situation.”
The extra logic circuitry in customized HBM means there may be extra alternative to make use of on-die screens on this base die to assist detect timing margin issues. “We take a look at the SoCs and HBMs as a system as a result of there’s interplay between the 2,” mentioned proteanTecs’ Brousard. “As an illustration, a variety of visitors coming in from the HBM may trigger a present surge, which results in a voltage droop. This can readjust, however that sudden voltage droop may cause failures, and a sensor with quick response time can seize that change. We wish to have that type of visibility, as a result of the SoC is being affected by the HBM. So it’s actually the work of many screens collectively that present the wealthy dataset wanted not solely to establish issues, however to deduce the rationale behind it in order that it may be correctly mitigated.”
HBM faces limits of probe-ability
Interconnect bump pitch between DRAMs in HBM is now so tight (<40µm, with 20 to 25µm microbumps), that it has turn into almost not possible to probe the microbumps straight. Even when the bumps may very well be probed, the possibilities for harm are too excessive. So in lots of instances, bigger sacrificial pads are used for probing, however in the long term it seems that engineers can be more and more depending on built-in self-test (BiST) choices, embedded screens and sensors, and redundancy and restore mechanisms to make sure greater interconnect yield in manufacturing and in area use.
“Any sign integrity points after meeting and multi-die packaging turn into harder to diagnose and debug since probing just isn’t simple,” mentioned Goriawalla. “Along with the lane restore capabilities constructed into the HBM protocol itself, Synopsys presents SLM ext-RAM IP that gives built-in redundancy evaluation and post-package restore to check the hundreds upon hundreds of interconnects, in HBM. Throughout any service downtime mode, the person can run JEDEC-recommended algorithms by way of SLM ext-RAM to carry out analysis and restore. This additionally permits a proactive method. As an illustration, you might understand that sure failures happen within the area as a consequence of phenomena like getting old or maybe some marginalities.You don’t need your LLM, which might take days or perhaps weeks to run, to fail due to this marginality. So that you proactively swap out a marginal lane with a superb lane, enabling that in-field, in-system.”
In-system testing and in-field analysis are extraordinarily engaging to hyperscalar clients which are always pushing their techniques for the very best uptime doable. “Throughout mission mode testing you might be figuring out whether or not the IC is performing its operate, as anticipated,” mentioned Goriawalla. “Because it does the inference or the coaching, you utilize a contactless embedded monitoring system to see amongst all these lanes within the for die-to-die interfaces, if, for instance, sure marginalities are occurring, or is the attention of the PHY getting smaller? Perhaps you could have a ‘strolling wounded’ interconnect, and as an alternative of letting it go to failure, on the subsequent scheduled downtime, you’re taking that marginal lane offline and swap it for a superb lane.”
In-system exams use embedded deterministic check (EDT) patterns to allow focused in-field testing to detect latent defects which will come up throughout any stage of the machine lifecycle. Throughout use, chip producers more and more have to account for the consequences of thermal stress, workload-induced degradation, voltage fluctuations, and different causes of getting old. For that motive, in-system check, as soon as restricted to automotive and mission-critical techniques, has discovered its method into knowledge facilities.
Bundle-level interface reliability
Verifying the connectivity and performance of XPU-XPU and XPU-HBM interfaces is determined by well-established boundary scan testing, reminiscence BiST, and at-speed purposeful testing. “Boundary scan, significantly 1149.1 and 1149.6, serves as a workhorse for testing interconnects at each the board and bundle ranges. Every chip, together with the xPU and HBM, incorporates a boundary scan register round its I/O pins. This methodology is efficacious for detecting opens, shorts, and stuck-at faults on the interconnects with no need to completely function the core logic,” mentioned Siemens EDA’s Phan. “For top-speed, differential AC-coupled interfaces, frequent in xPU-xPU hyperlinks, the 1149.6 commonplace is particularly designed to check their integrity, together with detecting shorts between differential pairs.
Phan famous that for HBM, devoted MBiST or customized MBiST may be applied throughout the SoC die. This MBIST generates particular knowledge patterns, drives them throughout the HBM interface, after which reads them again from the HBM dies, thereby verifying the complete knowledge path, together with interposer traces, microbumps, and HBM I/O logic. The HBM dies themselves typically function a loopback check mode that customers can make the most of to construct a BiST for die-to-die interconnect exams at-speed. HBM additionally helps lane restore functionality by the usual 1500 interfaces when a defective lane is detected.
“For top-speed serial hyperlinks between xPUs (like PCIe, CXL, or proprietary interconnects), specialised SerDes (serializer/deserializer) BiST is frequent. This entails activating loopback check modes (both internally or externally by way of bundle/interposer traces), and pseudo-random binary sequence (PRBS) era and checking,” mentioned Phan.
Purposeful at-speed exams are important when verifying connectivity. “These exams contain initiating giant knowledge transfers between the xPUs and HBM, or between two xPUs, and meticulously verifying the integrity of the information,” Phan mentioned. “This method exams the complete communication stack, encompassing protocols and error correction mechanisms, making certain that the interfaces carry out as anticipated underneath real-world knowledge hundreds.”
On the identical time, Phan emphasised the growing significance of I/O or lane restore capabilities. “These options forestall the necessity to discard a complete chip or bundle as a consequence of localized defects,” he mentioned. “This built-in redundancy is crucial for sustaining sign integrity and decreasing manufacturing waste in advanced AI accelerator techniques.”
Lengthy-term reliability and silicon lifecycle administration are additionally vital concerns in 2.5D/3D chiplet-based packages. “Growing old and stress-induced degradation require nearer coordination between {hardware} and system software program, with steady monitoring of parameters similar to delay shifts, course of monitoring, temperature, frequency, and eye width to allow well timed selections all through the chip’s lifecycle,” mentioned Surbhi Bansal, engineering director of DFT at Cadence. “In consequence, DFT is evolving past conventional structural check. It now incorporates cross-die observability, similar to in-situ eye-width measurement from the PHY throughout calibration and analog monitoring by ADC-based ATEST paths with digital readout, together with stress-aware check patterns. As well as, built-in redundancy and restore mechanisms throughout dies are a part of the JEDEC spec and supported by our HBM PHY. Collectively, these capabilities allow extra complete visibility and resilience throughout the complete lifecycle of the HBM stack.”
Failures in 2.5D/3D architectures, TSV defectivity
As HBM stacks develop from 8 chiplets to 12, 16, and past, issues that after had been mere nuisances have gotten main challenges. The pains of meeting result in warpage points, which precipitate in cracks or misalignment of interconnects. “Points similar to cracks (together with latent defects), thermal distribution, and hotspots have already proved vital, and these results could turn into much more pronounced with greater stack density,” mentioned Advantest’s Yokoyama. By way of DFT, “approaches similar to programmable MBiST with extra advanced inner sample execution could also be thought-about to enhance protection.”
A serious supply of defectivity is within the through-silicon vias (TSVs). With every new era of HBM, TSVs turn into extra intently spaced, introducing extra potential for failure. TSVs are lined with a skinny dielectric barrier earlier than the copper is deposited contained in the vias, and any discontinuities on this layer can show detrimental to yield. Because the pitch of TSVs shrinks, so does the pitch of the underlying microbumps that join every DRAM to the DRAM under it within the stack.
HBM producers sometimes use the by way of center integration move for TSV formation, that means the TSVs are shaped after the front-end processes (transistors), however earlier than back-end metallization. The benefit to this method is that the TSV may be built-in earlier than the wafer is thinned, whereas it’s nonetheless thick and mechanically strong. Forming TSVs at this juncture additionally avoids exposing the completed BEOL stack to the aggressive by way of etch, high-temperature oxide line deposition, and copper annealing steps. TSVs inside HBM are typically 2 to five microns in diameter and 30 to 60 microns deep (i.e., the depth of the wafer).
When bigger TSVs are uncovered on the wafer floor, they might be contacted by a probe. Nevertheless, given the tiny dimension of the TSV in HBM relative to probe needles (round 35µm), solely TSV check constructions may be contacted. Sometimes, lots of or hundreds of vias are linked in a daisy chain, significantly after bottom wafer thinning and TSV reveal, the place defects may be launched. Any deviation from a standard in-series resistance can point out potential defects similar to opens, cracks, incomplete metallic fill, misaligned contact, and so on. However an outlier daisy chain end result alone won’t spotlight which TSV has failed.
“With the normal daisy chain, they chain many bonds collectively, so actually you could have a continuity examine, a go/no go check. There’s a lack of details about the person bond as a result of it’s drowned out within the noise of the complete chain,” mentioned Modus Take a look at’s Lewis. “What we have to do within the check automobile design, as an alternative of simply chaining giant chains, is make a Kelvin connection and rise up round every certainly one of these bonds and measure the bond — every particular person bond with a Kelvin connection — very exactly into the micro-ohm vary, by distributing [test structures] throughout every layer of the substrate, every die-to-die interconnect. It’s all about planning, check insertion of the circuit, after which making the measurements and getting adequate knowledge for it to be priceless. We want hundreds of measurements per bundle to take that info and prepare the inspection fashions, as properly.”
A part of the change in TSV testing has to do with growing sign integrity (SI) and energy integrity (PI) points as TSVs get nearer collectively. “The business’s shift in TSV testing is being pushed by a number of elements,” mentioned Cadence’s Bansal. “Increased speeds end in SI/PI dominating performance. There’s additionally a necessity for at-speed, margin-aware validation, software program strategies to plot the attention, and system-level at-speed loopback exams.” She famous the necessity for extra complete LFSR polynomials to ship extra exhaustive seeds for testing SI/PI and different defects inside and throughout lanes.
When defects are discovered, HBM depends closely on redundancy and restore to enhance yield. Testing of the reminiscence array sometimes identifies defective rows or columns, spare sources are recognized, and the restore program is executed. “Reminiscence restore is a built-in function of the HBM die itself, as a result of every die comprises reminiscence BiST engines particularly tailor-made to run refined reminiscence check algorithms to completely check the reminiscence cells,” mentioned Phan. “As soon as a defective reminiscence location is recognized (by the HBM’s BiST or the SoC’s MBiST), the person can program the failing addresses and channels into SOFT_REPAIR or HARD_REPAIR Write Information Registers (WDRs) by way of the IEEE-1500 interfaces. After programming, a reminiscence restore operation may be initiated, permitting the HBM to reconfigure itself to bypass the defective components, bettering yield and reliability.”
There’s an growing want for power-aware simulation due to the nice variety of high-speed channels working concurrently. “Energy integrity is a rising problem,” mentioned Bansal. “The growing variety of high-speed channels introduces better susceptibility to IR drop and simultaneous switching noise, driving the necessity for extra strong design and power-aware simulation methods, similar to staggered activation to scale back preliminary inrush present.”
One other expertise that may assistance is power-aware automated check program era. “Energy-aware ATPG is commonly a necessity for AI purposes as a consequence of their high-power calls for,” mentioned Siemens EDA’s Phan. “It might reduce energy consumption in the course of the crucial scan shift and scan seize operations by a mix of {hardware} insertion and sample era methods.” When mixed with the Streaming Scan Community (SSN), power-aware ATPG can clean the ability profile by staggered shift clocks, decreasing check time and energy utilization considerably.
Excessive temperature throughout testing additionally contributes to failures. “AI workloads make the most of full HBM bandwidth on a steady foundation. Therefore, warmth era is turning into more and more frequent with HBM,” mentioned Bansal. “As an illustration, excessive toggle charges trigger localized heating, leading to timing shifts, IR drop, and elevated leakage. Throughout check, warmth era is a matter with a excessive ATPG toggle fee. Nevertheless, this could additionally assist display the components throughout check.”
Collectively, these failure modes name for silicon lifecycle monitoring by manufacturing and into the sphere to allow software program restore at any time when doable and {hardware} replacements solely as wanted. Typically, there’s a development to make use of knowledge middle techniques longer, which can solely be doable when satisfactory in-field check and restore mechanisms are on board.
Conclusion
As HBMs develop in stack top and complexity, so does the testing method and design-for-test technique. The business is within the technique of rolling out HBM4 gadgets with stack heights of 16 DRAMs, a few of which may have customized logic base dies to customise the configuration for particular workloads. DFT too can be custom-made to satisfy specialised wants and tackle the various failure modes of stacked die, together with TSV movie discontinuities and unlanded contacts, thermally-induced failures, timing shifts, SI/PI failures, die cracks, and so on.
An vital function of DFT is to confirm the connectivity and performance of xPU-xPU and xPU-HBM interfaces utilizing boundary scan testing, reminiscence BiST, and at-speed purposeful testing. Embedded screens zoom in on crucial metrics like timing margin, temperature adjustments, and voltage droop, enabling huge knowledge evaluation and potential tracing of failures to their root trigger. Extremely-precise Kelvin measurements in applicable check constructions might help guarantee the standard of particular person bonds, whether or not they’re hybrid-bonded chip-to-chip or thermocompression-bonded microbumps.
What’s clear is {that a} plethora of instruments and DFT methods are wanted to completely check superior interconnects in stacked chips shifting ahead, and HBM is the proving floor for these strategies.
Associated Articles
Multi-Die Testing In The Field Must Build On Established Test Methodologies
Having a tool that works at time zero is now not a assure of reliability over its lifetime.
AI Accelerators Usher In New Era For IC Test
The quantity and number of check interfaces, coupled with elevated packaging complexity, are including a slew of recent challenges.
AI Accelerator Testing Depends On DFT Innovations
Multi-die assemblies vastly enhance the variety of issues that may go fallacious, and the issue of discovering them.
HBM Shifts Testing Left To Preserve AI Chip Yield
Testing sooner and extra typically can enhance high quality and scale back scrap, nevertheless it’s additionally extra pricey.