Skip to content

About

Silicon as a Distributed System: Closing the Cross-Layer Loop from Spatial Manufacturing Defects to Wafer-Level Economics

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

Silicon as a Distributed System: Closing the Cross-Layer Loop from Spatial Manufacturing Defects to Wafer-Level Economics

Paper PDF License: MIT Python 3.10+

Author: Ishwar Chand Meena
Indian Institute of Technology Bhilai
📄 Preprint: paper/silicon_as_distributed_system.pdf


Overview

As semiconductor scaling approaches reticle limits and deep learning accelerators expand beyond $500,\text{mm}^2$, zero-defect manufacturing yields collapse exponentially. Traditional semiconductor manufacturing treats partially defective dies as binary scrap or relies on rigid hardware spare rows with limited spatial salvage potential. In contrast, distributed computing systems routinely deliver reliable, deterministic computation across unreliable, heterogeneous nodes through software-managed virtualization, message-passing isolation, and dynamic task rescheduling.

This repository provides an end-to-end evaluation framework treating silicon as a distributed system. We close the quantitative causal loop from physical wafer-scale defect distributions to cycle-accurate on-chip network contention, deterministic multi-pass workload completion, qualified-die yield, and commercial economic operating boundaries.

Calibrated against open foundry standard-cell PDKs (SkyWater SKY130 and GlobalFoundries GF180MCU), the model couples negative-binomial spatial defect clustering and fatal uncore periphery failures to a cycle-accurate flit-level 2D mesh/torus network simulator executing a Transformer Attention Projection GEMM ($M=K=N=2048$, $8.59 \times 10^9$ MACs). We demonstrate that software-controlled multi-pass time-multiplexing and deadlock-free adaptive detour routing convert spatial physical defects into deterministic temporal performance degradation, achieving 100% mathematical completion for evaluated workloads on qualified dies.


Research Questions & Headline Findings

RQ1: NoC Contention & Throughput Saturation

  • Finding: Under West-First deadlock-free detour routing, aggregate delivered throughput derates gracefully from $22.52,\text{flits/cycle}$ (pristine $8 \times 8$ mesh) to $19.12,\text{flits/cycle}$ at 5% faults, $14.12,\text{flits/cycle}$ at 10% faults, and $12.44,\text{flits/cycle}$ at 20% faults. Average path stretch remains bounded between $1.356\times$ and $1.439\times$.

RQ2: Wafer-Scale Yield Recovery & Complexity Scaling

  • Finding: On a standardized $16 \times 16$ ($5.76,\text{cm}^2 = 576.0,\text{mm}^2$) accelerator die under an externally motivated 9.6% Vicis silicon area tax anchor:
    • At $D_0 = 0.25,\text{def/cm}^2$: Rigid yield is 49.8%, adaptive qualification achieves 91.8% yield (salvage fraction $R_{\text{salvage}} = \mathbf{83.6%}$, recovering $+33.3$ net usable dies/wafer).
    • At $D_0 = 0.50,\text{def/cm}^2$: Rigid yield drops to 25.6%, adaptive qualification maintains 85.2% yield ($R_{\text{salvage}} = \mathbf{80.1%}$, recovering $+50.9$ net usable dies/wafer).
    • At $D_0 = 1.00,\text{def/cm}^2$: Rigid yield collapses to 8.18%, adaptive qualification preserves 72.5% yield ($R_{\text{salvage}} = \mathbf{70.1%}$, recovering $+56.6$ net usable dies/wafer).

RQ3: Adversarial Robustness & Break-Even Boundary

  • Finding: Defect tolerance does not universally outperform traditional qualification. Rather, it embodies an architectural trade-off with an identifiable economic operating envelope:
    • Model 1 ($M_1$, Strict Integrity): Break-even defect density $D_{\text{break-even}} \in [\mathbf{0.124, 0.140}],\text{defects/cm}^2$.
    • Model 2 ($M_2$, Sub-SKU Down-Binning): $D_{\text{break-even}} \in [\mathbf{0.151, 0.170}],\text{defects/cm}^2$.
    • Model 3 ($M_3$, Continuous TFLOPS Pricing): $D_{\text{break-even}} \in [\mathbf{0.208, 0.236}],\text{defects/cm}^2$.
    • Regime Boundary: When silicon is nearly flawless ($D_0 < 0.12,\text{cm}^{-2}$), paying an upfront area tax destroys good dies via Gross Dies per Wafer (GDW) shrinkage. Inside the salvage window ($0.12 \le D_0 \le 1.00,\text{cm}^{-2}$), adaptive qualification delivers substantial profit deltas ($+$3.4\text{k}$ to $+$16.2\text{k}$ per wafer).

The Cross-Layer Framework

   ┌────────────────────────────────────────────────────────┐
   │ 1. Open PDK Physical Grounding (SKY130 / GF180MCU)     │
   │    Standard-cell areas, M1 RC delay, calibrated Fmax   │
   └───────────────────────────┬────────────────────────────┘
                               ▼
   ┌────────────────────────────────────────────────────────┐
   │ 2. Spatial Defect & Uncore Fatality Generator          │
   │    Negative-binomial compound Poisson (alpha=2.0)      │
   └───────────────────────────┬────────────────────────────┘
                               ▼
   ┌────────────────────────────────────────────────────────┐
   │ 3. Cycle-Accurate Flit NoC (4 VCs, West-First Detours) │
   │    Synthetic injection sweeps & saturation curves      │
   └───────────────────────────┬────────────────────────────┘
                               ▼
   ┌────────────────────────────────────────────────────────┐
   │ 4. Deterministic Multi-Pass AI Workload Remapper       │
   │    Transformer Attention Projection GEMM (2048^3)      │
   └───────────────────────────┬────────────────────────────┘
                               ▼
   ┌────────────────────────────────────────────────────────┐
   │ 5. Wafer-Scale Monte Carlo Economic Boundary Engine    │
   │    Models M1 (Integrity), M2 (Binning), M3 (Throughput)│
   └────────────────────────────────────────────────────────┘

Repository Structure

silicon-as-distributed-system/
├── README.md                          # This document
├── LICENSE                            # MIT License
├── requirements.txt                   # Python dependencies
│
├── paper/                             # Publication artifacts
│   ├── silicon_as_distributed_system.pdf  # Compiled 9-page preprint
│   ├── main.tex                       # LaTeX source code
│   ├── references.bib                 # 13 verified BibTeX citations
│   └── IEEEtran.bst                   # IEEE bibliography style
│
├── figures/                           # Publication-grade evaluation figures
│   ├── fig1_wafer_defects.png         # Wafer defect map under compound Poisson clustering
│   ├── fig2_die_remapping.png         # Spatial fault maps, allocation schedules, detours
│   ├── fig3_noc_saturation.png        # Latency vs. injection rate under 0%-20% faults
│   ├── fig4_gemm_benchmark.png        # GEMM makespan & multi-pass serialization
│   ├── fig5_thermal_hotspots.png      # 2D steady-state thermal hotspot profile
│   ├── fig6_stress_tested_crossover.png # 16x16 monolithic stress test crossover
│   ├── fig7_uncore_sensitivity.png    # Uncore periphery area sweep (0%-25%)
│   └── fig8_adversarial_heatmaps.png  # Multi-model profit delta heatmaps
│
└── experiments/                       # Complete simulation & analysis engine
    ├── EXPERIMENT_PROVENANCE_LEDGER.md # Ledger mapping claims to raw data
    ├── open_pdk_calibrator/           # Layer 1: PDK cell extraction & tile physics
    ├── cycle_accurate_noc/            # Layer 2: Flit-level NoC simulation
    ├── real_ai_workload_remapper/     # Layer 3: GEMM trace generation & remapping
    ├── wafer_scale_monte_carlo/       # Layer 4: Wafer Monte Carlo & economics
    ├── stress_tests/                  # Monolithic 16x16 stress test & thermals
    └── data/                          # Audited raw experiment CSVs

Reproduction Instructions

1. Environment Setup

git clone https://github.com/ishwar170695/silicon-as-distributed-system.git
cd silicon-as-distributed-system
pip install -r requirements.txt

2. Run Experiments

Step 1: PDK Standard-Cell Calibration

Grounds standard-cell dimensions, wire parasitics, and operational frequencies in open foundry PDKs (SkyWater SKY130 and GlobalFoundries GF180MCU):

python experiments/open_pdk_calibrator/physical_model.py

Step 2: Cycle-Accurate NoC Saturation Sweep

Evaluates flit-level packet transport and West-First detour routing across fault patterns:

python experiments/cycle_accurate_noc/saturation_sweep.py

Step 3: Systolic GEMM Workload Remapping

Simulates the $2048^3$ Transformer Attention Projection workload and multi-pass serialization:

python experiments/real_ai_workload_remapper/run_gemm_benchmark.py

Step 4: Wafer-Scale Monte Carlo Yield & Scaling

Executes the standardized 30-wafer Monte Carlo campaign with cluster-resampled 95% confidence intervals:

python experiments/wafer_scale_monte_carlo/complexity_scaling_experiment.py

Step 5: Adversarial Economic Sweeps

Sweeps 125 combinations of area tax, qualification threshold, and defect density across market models $M_1, M_2, M_3$:

python experiments/wafer_scale_monte_carlo/adversarial_robustness_grid.py
python experiments/wafer_scale_monte_carlo/uncore_sensitivity_sweep.py

Step 6: Monolithic $16 \times 16$ Stress Test & Thermal Modeling

Executes the 256-tile monolithic die evaluation with 2D steady-state Fourier conduction and DVFS derating:

python experiments/stress_tests/stress_test_experiment.py

Threats to Validity & Engineering Limitations

  1. Binary Defect Assumption: Simulation treats defects as binary, diagnosed, and cleanly isolated. Latent defects (gate oxide pinholes, high-resistance vias) cause test escapes that software remapping cannot prevent.
  2. Buffer Memory Overhead: Multi-pass execution assumes intermediate activation tensors fit in on-chip SRAM/scratchpads. Spilling partial sums to off-chip HBM introduces memory bandwidth bottlenecks.
  3. Interconnect Deadlock Scope: West-First routing provably eliminates resource cycles on 2D meshes, but 2D torus wrap-around channels introduce cyclical dependencies that require virtual channel partitioning for general deadlock freedom.
  4. Compiler NRE: Generating unique per-die binary configurations, placement embeddings, and routing tables imposes software engineering costs that must be amortized over high-volume production.
  5. Thermal Proxy Simplification: The thermal model is a parameterized 2D steady-state Fourier-conduction proxy, capturing localized heating gradients without substituting for transient 3D CFD modeling.

Citation

@article{meena2026silicon,
  author    = {Ishwar Chand Meena},
  title     = {Silicon as a Distributed System: Closing the Cross-Layer Loop from Spatial Manufacturing Defects to Wafer-Level Economics},
  year      = {2026},
  publisher = {GitHub},
  journal   = {GitHub repository},
  howpublished = {\url{https://github.com/ishwar170695/silicon-as-distributed-system}}
}

License

This project is licensed under the MIT License — see the LICENSE file for details.

About

Silicon as a Distributed System: Closing the Cross-Layer Loop from Spatial Manufacturing Defects to Wafer-Level Economics

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages