Manufacturing Quality Control
The capstone use case: a real process engineer’s workflow on real UCI SECOM semiconductor manufacturing telemetry: 1,567 real production runs, 590 real anonymized sensors, 104 real failures (6.6%), genuinely messy (116 dead sensors, 538 with real missing readings).
Feature triage, honestly documented
Section titled “Feature triage, honestly documented”Drop dead and >50%-missing sensors, rank survivors by |correlation| with the real
fail label, greedily decorrelate (skip anything >0.9 correlated with an
already-picked sensor). A naive top-15-by-correlation pick turns out to produce a
covariance matrix within 1e-14 of singular, a real numerical trap this triage
avoids.
Hotelling’s T² via real INVERSE
Section titled “Hotelling’s T² via real INVERSE”Standardizing against the pass-only reference distribution turns a near-singular raw
covariance (det ~ 7e-31, condition number ~8e12) into a well-behaved one
(det ~ 0.11, condition ~10), a textbook illustration of why standardizing before
a covariance-based method isn’t just convention:
LET XcT = TRANSPOSE Xc_pass_ds_arrayLET cov_raw = MATMUL XcT Xc_pass_ds_arrayLET cov = SCALE cov_raw BY 0.000684LET cov_det = DETERMINANT covLET cov_inv = INVERSE covDETERMINANT/INVERSE matched numpy to 3.00e-05 max difference. Flagging the top
200 of 1,567 runs by T² score catches 20/104 real fails, a modest, honestly-reported
1.5x lift over baseline, not oversold.
Virtual metrology via SOLVE/QR/LU
Section titled “Virtual metrology via SOLVE/QR/LU”Predicting one sensor from 14 others (a real fab technique for expensive-to-measure
signals), solved via the normal equations, with QR/LU used as real
decomposition-correctness checks rather than a second way to derive the same answer:
LET beta = SOLVE XtX XtyLET q, r = QR XtX -- Q@R reconstructs X^T X; Q^T@Q ~= ILET p, l, u = LU XtX -- P @ X^T X reconstructs L @ UR² = 0.278 predicting one sensor from the others, modest and honest, these are
weakly-correlated fab sensors.
Named, chained pipelines: first time in this project
Section titled “Named, chained pipelines: first time in this project”DATASET runs COLUMNS (run_id: Int, batch_id: String, label: Int, t2_score: Float)DEFINE PIPELINE flagged AS WHERE t2_score > 10.0 THEN ORDER BY t2_score DESCAPPLY PIPELINE flagged ON runs INTO stage_flaggedSELECT * FROM stage_flaggedTwo pipelines chained (the second’s input is the first’s output), mirroring a real
analyst’s “filter, then rank the survivors” workflow, then round-tripped through
SAVE/LOAD PIPELINE.
Relational analytics untested before this notebook
Section titled “Relational analytics untested before this notebook”HAVING on a real aggregate condition, LAG/LEAD for trend detection,
COALESCE/NULLIF on genuinely missing sensor readings: all new DSL surface for
this project:
SELECT batch_id, COUNT(*) AS n, AVG(label) AS fail_rate FROM runsGROUP BY batch_id HAVING AVG(label) > 0.15 AND COUNT(*) >= 5Governance and a real operational bonus
Section titled “Governance and a real operational bonus”ATTACH/AUDIT DATASET (referential-integrity governance), EXPLAIN LINEAGE, and
EXPORT, including a dataset with a real Vector(15) column, which used to crash
the CSV writer outright. In a clearly separated bonus section: the same unmodified DSL
submitted as a real background job and registered as a real recurring /schedule
task against an actual linal serve process, closing the “can this run unattended?”
question directly.
Two real engine bugs, found and fixed here
Section titled “Two real engine bugs, found and fixed here”LET <new_name> = <existing_tensor_name>(bare-identifier alias) silently failed to bind the new name: no error, and the success message even named the old variable.EXPORTing a dataset with aVector/Matrixcolumn to CSV crashed outright.
Both fixed and shipped same-session, in PR
#101 and PyPI linaldb
0.1.8; this notebook, re-run against
that real published wheel, demonstrates both fixes directly with zero workarounds.

