Search papers, labs, and topics across Lattice.
This paper introduces an exact multistate reliability framework for High Bandwidth Memory (HBM) systems, allowing for the modeling of service units that can deliver varying levels of bandwidth rather than just binary outcomes. The proposed threshold-pruned multistate Binary-Addition-Tree (TP-mBAT) algorithm significantly reduces computational complexity while maintaining accuracy, achieving a drastic reduction in peak storage requirements. Key findings reveal that neglecting shared stress among service units can lead to an overestimation of reliability by nearly 9 percentage points, emphasizing the importance of considering multistate interactions in HBM design.
Ignoring shared stress in HBM systems can inflate reliability estimates by nearly 9 percentage points, revealing critical insights for system design.
High Bandwidth Memory (HBM) systems can exhibit partial service rather than only full service or complete isolation: a controller-visible service unit may deliver full, reduced, or zero bandwidth because of sub-channel isolation, lane remapping, or protection overhead. This paper develops an exact multistate reliability framework in which each service unit carries an arbitrary finite set of bandwidth states and reliability is the probability that aggregate delivered bandwidth meets a demand. The binary k-out-of-n model is recovered as a special case, while closed-form binary-mapping error relations quantify mean-bandwidth distortion and provide a screening test for the simpler abstraction. For exact single-threshold evaluation a threshold-pruned multistate Binary-Addition-Tree (TP-mBAT) algorithm is proposed. It is deliberately regime-specific: fixed-grid dynamic programming is preferable on a compact common grid, where a 16-unit commensurate control required 273 pruned-DP updates versus 362,506 TP-mBAT node visits. On a reproducible 14-unit incommensurate benchmark, TP-mBAT is compared against a dynamic program carrying the same threshold rules, so that no baseline is weakened. Both expand the same state space, 6,862 nodes against 6,861 updates, and the separation lies entirely in retained state: 17 traversal entries against 412,121 probability states at central demand, reducing measured peak storage from 25.02 MB to 1,152 B. An exact probability-transfer sensitivity identifies when moving mass from a degraded state to a higher-bandwidth state changes system success, and a reserved-unit floor model admits a third exact pruning rule that is vacuous without such floors. A latent package-state mixture captures shared stress, where ignoring dependence overstates reliability by 8.73 percentage points.