Module 4 · Measure
Measurement systems analysis
Before Measure collects a single production value, it has to answer a question that has nothing to do with the process: can the gauge be trusted? Every number in this course, from a mean to a Cpk to a p-value, is a number the gauge reported, not a number the process actually produced. If the gauge's own variation is a large share of what gets measured, no amount of statistical sophistication downstream can tell the difference between a process problem and a measurement problem. This module covers the full measurement systems analysis (MSA) toolkit: bias, linearity and stability studies, repeatability and reproducibility, crossed gauge R&R by the ANOVA method, attribute agreement analysis, and what to actually do when a gauge fails.
Learning objectives
- Explain why a measurement system must be verified before a process capability study is meaningful.
- Run a bias and linearity study against certified reference standards and test whether the bias is statistically significant.
- Distinguish repeatability from reproducibility and compute both from a crossed gauge R&R study.
- Interpret a crossed gauge R&R ANOVA table: variance components, %Contribution, %Study Variation, %Tolerance, and ndc.
- Run an attribute agreement study and interpret Cohen's and Fleiss' kappa.
- Decide what to do when a gauge study fails, and know which fixes address which failure.
Why this matters
NIST's e-Handbook frames the problem in one line: a measurement is not the true value of a characteristic, it is the true value plus measurement error, and that error has its own repeatability, reproducibility and stability that must be characterized before the measurement can be trusted for anything else.[1] Burdick, Borror and Montgomery's 2003 review of gauge capability methods makes the same point from the ANOVA side: the variance an inspector observes in a set of measurements is not the part-to-part variance alone, it is part variance plus operator variance plus a part-by-operator interaction plus pure measurement error, all superimposed, and a two-factor ANOVA is the standard way to pull them apart.[2] Total observed variance decomposes as σ²observed = σ²process + σ²measurement. A process with genuinely tight, capable production can still produce a terrible-looking Cpk if the gauge measuring it is noisy; a genuinely loose process can look artificially capable if the gauge is too coarse to resolve the scatter. Module 3's crankshaft case is a reminder of the same discipline from the other direction: a DMAIC project that reports a Cp or a Cpk owes the reader confidence that the number describes the process and not the gauge.
The scenario below is a constructed illustration, not a real incident.
A common, entirely avoidable failure: a machining cell gets chartered for a capability problem because its bore diameter Cpk has been drifting for three months. Two weeks into Measure, someone finally runs a gauge R&R on the bore gauge that has been reporting every one of those numbers, and finds %GRR of 45 % of the tolerance — the gauge itself, not the machining process, explains most of the apparent variation. The three months of capability trend data, every control chart built from it, and the machine adjustments made in response to "special causes" that were actually gauge noise, all have to be set aside. The fix (a worn gauge anvil, replaced for the cost of a spare part) takes an afternoon. The MSA study that should have caught it in week one, run last, cost the project most of a quarter. Running the gauge study first is not bureaucracy; it is the cheapest possible insurance against exactly this outcome.
Why the gauge comes before the process
Every characteristic this course measures — a bore diameter, a pull strength, a leak-test result — is observed through some instrument: a micrometer, a load cell, a go/no-go fixture, a trained eye. The observed value is never the part's true dimension; it is the true dimension plus whatever the gauge and the person using it add or subtract on that particular reading. If that added variation is small relative to the process's own spread, it can be ignored. If it is not small, every downstream number is compromised: a capability index computed from noisy data understates true capability, a control chart built on a noisy gauge has wider limits than the process itself would produce and so misses real shifts, and a hypothesis test comparing two conditions loses power because gauge noise is added to both sides. MSA answers one question before any of that work starts: how much of what I am about to measure is the part, and how much is the measuring?
Bias, linearity, and stability
These three studies ask whether the gauge, on average, reads the true value (bias), whether that accuracy holds across the whole range the gauge is used over (linearity), and whether it stays that way over time (stability). All three need one thing the process itself does not require: a certified reference standard with a known true value.
Bias
A small bias is not automatically a problem; a bias that is both statistically significant and large relative to the tolerance is. AIAG's guidance, echoed across the MSA literature, is to report bias as a percentage of the process tolerance so an engineer can judge practical significance alongside the t-test's statistical significance.
Linearity
A gauge can be unbiased on average across its whole range and still be biased low at one end and high at the other; averaging those two errors together can hide a real problem. A linearity study measures several certified reference standards spread across the gauge's operating range, computes the bias at each level, and regresses bias against the reference value. A flat line at bias = 0 is the ideal; a non-zero, statistically significant slope means the gauge's error itself depends on where in the range it is reading.
Stability
Bias and linearity are snapshots. Stability asks whether that snapshot holds up over weeks of use: does the same reference standard, measured once a shift, drift as the gauge wears, gets recalibrated, or is handled by different people over time? An I-MR chart (Module 16 builds the general tool; here it is applied to gauge readings instead of process output) of repeated reference readings answers this the same way a process control chart answers it for a process: points beyond the 3-sigma limits, or a run, mean something changed.
Repeatability and reproducibility
Repeatability is the variation in repeated measurements of the same part, by the same appraiser, with the same gauge, close together in time — the gauge's own inherent noise, sometimes called equipment variation (EV). Reproducibility is the additional variation that shows up when different appraisers measure the same parts with the same gauge — operator-to-operator variation, sometimes called appraiser variation (AV). A gauge can be highly repeatable (an individual operator gets nearly identical readings every time) and still poorly reproducible (three different operators, each internally consistent, systematically disagree with each other), and a study has to check both, because they call for different fixes: a repeatability problem is usually the instrument or the fixture, a reproducibility problem is usually training, ergonomics, or an ambiguous operational definition (Module 5 covers operational definitions directly).
Crossed gauge R&R by the ANOVA method
A crossed study has every appraiser measure every part the same number of times, in randomized order, usually blind to their own prior readings. "Crossed" means every appraiser-part combination is observed, which is what lets a two-way ANOVA separate part variation, operator variation, and the part-by-operator interaction (does a particular operator read particular parts differently than other operators do) from pure repeatability error, following Burdick, Borror and Montgomery's review of the method.[2]
Reproducibility (appraiser variation) = the operator variance component + the interaction variance component
Total gauge R&R variance = repeatability variance + reproducibility variance Variance components are recovered from the ANOVA mean squares by subtraction (each mean square's expected value under the random-effects model minus the next one down, divided by the appropriate replication count), the same logic NIST uses for nested gauge studies.[3] A first F-test on the interaction term decides whether to keep it or pool it into repeatability; AIAG and Minitab both default to pooling when that test's p-value exceeds 0.05.
%Study Variation = 100 × (standard deviation / total standard deviation), using k = 6 (99.73 % of a normal distribution) or k = 5.15 (99 %) times the standard deviation as "study variation"
%Tolerance = 100 × (k × standard deviation / tolerance)
ndc (number of distinct categories) = floor(1.41 × σpart / σGRR) %Contribution uses variances (so the gauge and part shares sum to 100); %Study Variation and %Tolerance use standard deviations, so they do not sum to 100 the same way and can tell different stories about the same gauge, particularly when the part sample does not span the full tolerance. ndc estimates how many truly distinguishable groups the gauge can separate parts into — AIAG's rule of thumb is 5 or more.
AIAG's Measurement Systems Analysis reference manual sets out the acceptance bands used across the industry: %GRR (the larger of %Study Variation and %Tolerance) under 10 % is generally acceptable, 10 to 30 % may be acceptable depending on the application and the cost of a better gauge, and over 30 % is unacceptable.[4][8], [9] These bands, and the ndc ≥ 5 rule of thumb, are guidelines a customer can tighten or relax by agreement, not laws of statistics; a safety characteristic may reasonably demand better than 10 %, and a cosmetic one may reasonably tolerate more.
Attribute agreement analysis
Not every characteristic is measured on a continuous scale. A visual accept/reject call, a go/no-go gauge, a pass/fail leak test — all of these need their own kind of measurement system study, because "repeatability" for a discrete judgment is not a standard deviation, it is whether the same appraiser gives the same answer twice, and whether different appraisers agree with each other and, if a known correct answer exists, with the truth.
Cohen's kappa measures chance-corrected agreement between two raters on a categorical judgment.[5] Fleiss' kappa extends the idea to three or more raters.[6] Both start from the same idea: raw percent agreement overstates how good a rater is, because two raters who both call 90 % of parts "accept" will agree with each other about 82 % of the time by chance alone, even if neither one is looking at the part. Kappa subtracts out that chance agreement.
Landis and Koch's widely used benchmark scale reads κ ≤ 0 as poor, 0.01 to 0.20 as slight, 0.21 to 0.40 as fair, 0.41 to 0.60 as moderate, 0.61 to 0.80 as substantial, and 0.81 to 1.00 as almost perfect agreement.[7] Like the AIAG %GRR bands, this is a rule of thumb the literature has converged on, not a statistical law; a safety inspection may need "substantial" or better before the study is accepted at all.
What to do with a bad gauge
A failed study is a diagnosis, not a dead end, and which fix applies depends on which number failed.
- High repeatability variation (equipment/EV). Look at the instrument itself: resolution too coarse for the tolerance, a worn fixture or anvil, an unstable mounting, vibration. Often the cheapest fix: better fixturing, not a new gauge.
- High reproducibility variation (appraiser/AV), especially interaction. Look at training and the operational definition: do operators handle the part the same way, read the display the same way, and agree on what "the measurement point" actually is? Retraining and a written, specific operational definition (Module 5) fix this more often than new equipment does.
- Significant bias. Recalibrate against the reference standard, or apply a documented correction factor if recalibration is not immediately possible.
- Significant linearity. A single-point calibration will not fix this; the gauge needs correction (or replacement) across its working range, not just at one setpoint.
- Instability. Investigate what changed: a component wearing, a new calibration standard, a new operator population. Fix the cause, then re-run the stability study before trusting the gauge again.
- Low attribute agreement. Usually a definition problem before it is a people problem: agree, in writing and with reference photos or masters, exactly what "reject" looks like, then retrain and re-run the study.
What never fixes a bad measurement system: widening the acceptance criteria until the gauge passes, or quietly reporting the process capability anyway with a footnote. If the gauge cannot be trusted, no number computed from it can be trusted either, and that has to be the conclusion stated to the sponsor, however unwelcome.
Worked examples
Worked example 1: bias and linearity, the bore gauge
The data below is a constructed example, not a real production run.
The bore gauge used throughout Modules 7 and 16 for the Ø12.000 ± 0.025 mm running example gets its own MSA study here. First, bias: a certified reference standard with a known diameter of 12.010 mm is measured 15 times in one sitting.
| Reading | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mm | 12.012 | 12.011 | 12.012 | 12.013 | 12.015 | 12.013 | 12.012 | 12.012 | 12.014 | 12.015 | 12.013 | 12.011 | 12.011 | 12.015 | 12.013 |
0.0028 mm against a 0.050 mm tolerance is 5.6 % of tolerance — worth a calibration adjustment, not an emergency. Next, linearity: five certified reference standards spanning the tolerance, four repeat readings each.
| Reference (mm) | Measured (mm) | Bias (mm) |
|---|---|---|
| 11.985 | 11.984 | −0.001 |
| 11.985 | 11.985 | 0.000 |
| 11.985 | 11.984 | −0.001 |
| 11.985 | 11.982 | −0.003 |
| 11.995 | 11.996 | 0.001 |
| 11.995 | 11.996 | 0.001 |
| 11.995 | 11.994 | −0.001 |
| 11.995 | 11.996 | 0.001 |
| 12.005 | 12.007 | 0.002 |
| 12.005 | 12.007 | 0.002 |
| 12.005 | 12.007 | 0.002 |
| 12.005 | 12.007 | 0.002 |
| 12.015 | 12.017 | 0.002 |
| 12.015 | 12.018 | 0.003 |
| 12.015 | 12.017 | 0.002 |
| 12.015 | 12.019 | 0.004 |
| 12.025 | 12.029 | 0.004 |
| 12.025 | 12.029 | 0.004 |
| 12.025 | 12.028 | 0.003 |
| 12.025 | 12.029 | 0.004 |
Average bias runs from about −0.00125 mm at the low end (11.985 mm) to about 0.00375 mm at the high end (12.025 mm) — small numbers, but they move in one direction as the reference value increases, which is exactly what a linearity problem looks like.
Both findings point the same way: this gauge reads a little high, increasingly so toward the top of the range. A single-point recalibration at mid-range would fix the average bias but leave the linearity trend untouched; a multi-point calibration correction, or gauge maintenance across the full range, is the appropriate fix (see "What to do with a bad gauge" above).
Worked example 2: stability, the same reference standard over time
Constructed data, not a real production run.
The same 12.010 mm reference standard is measured once at the start of each of 24 shifts, to see whether the bias found above holds steady or drifts.
| Shift | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mm | 12.013 | 12.014 | 12.013 | 12.011 | 12.014 | 12.013 | 12.012 | 12.014 | 12.013 | 12.013 | 12.013 | 12.014 | 12.012 | 12.013 | 12.012 | 12.014 | 12.013 | 12.012 | 12.012 | 12.012 | 12.013 | 12.012 | 12.015 | 12.014 |
No point beyond the limits, no run, no trend: this gauge is stable over the 24 shifts studied, even though it carries the small bias and linearity effects found above. Stability and accuracy are different questions; a gauge can be perfectly consistent about being slightly wrong.
Worked example 3: crossed gauge R&R, two gauges compared
Constructed data, not a real production run.
Two crossed studies, same design (10 parts × 3 operators × 3 trials, randomized order), on two different gauges: the bore gauge above, and a checkweigher used for fill-weight inspection (spec 248.0 to 254.0 g, tolerance 6.0 g). Both ANOVA tables test the part × operator interaction first; in both cases the interaction is not significant (bore gauge p = 0.229; checkweigher p = 0.162) and is pooled into repeatability before the variance components are computed.
| Quantity | Bore gauge (Ø12.000 ± 0.025 mm) | Checkweigher (248.0-254.0 g) |
|---|---|---|
| %Contribution (GRR) | 5.7 % | 0.5 % |
| %Study Variation (GRR) | 23.9 % | 7.0 % |
| %Tolerance (GRR) | 29.8 % | 11.4 % |
| ndc | 5 | 20 |
| AIAG verdict | Conditional (10-30 % band) | Acceptable (under 10 %) |
Same study design, same analysis method, two very different verdicts. The bore gauge's %GRR sits in AIAG's conditional band on both the study-variation and tolerance figures, with ndc right at the minimum of 5 — usable, but a candidate for the fixturing and resolution review described above, especially given the bias and linearity findings on this same gauge in Worked example 1. The checkweigher is comfortably acceptable on every figure. This is also why Module 7's bore capability numbers (Cwk about 1.06) carry a caveat this module makes explicit: with %GRR near 24 % of study variation, some of the "process" variation that capability index reports on is genuinely the gauge's.
Gauge R&R (crossed, ANOVA method)
Pre-loaded with the bore gauge study above (10 parts × 3 operators × 3 trials). Paste your own measurements to run a different study.
Show the checkweigher study data (10 parts × 3 operators × 3 trials, 90 readings)
| Part | A1 | A2 | A3 | B1 | B2 | B3 | C1 | C2 | C3 |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 248.6 | 248.6 | 248.6 | 248.7 | 248.7 | 248.6 | 248.6 | 248.6 | 248.7 |
| 2 | 249.2 | 249.4 | 249.4 | 248.9 | 249.0 | 249.2 | 249.2 | 249.3 | 249.3 |
| 3 | 250.0 | 249.6 | 249.7 | 250.0 | 249.8 | 249.8 | 249.7 | 249.5 | 249.8 |
| 4 | 250.0 | 249.9 | 249.9 | 250.0 | 249.9 | 250.0 | 250.0 | 250.0 | 250.0 |
| 5 | 250.9 | 251.0 | 250.9 | 250.8 | 251.0 | 250.8 | 251.0 | 250.8 | 251.0 |
| 6 | 251.3 | 251.2 | 251.3 | 251.3 | 251.4 | 251.2 | 251.2 | 251.4 | 251.3 |
| 7 | 251.7 | 251.8 | 251.5 | 251.8 | 251.9 | 251.7 | 251.6 | 251.8 | 251.8 |
| 8 | 252.5 | 252.4 | 252.2 | 252.4 | 252.4 | 252.5 | 252.5 | 252.3 | 252.3 |
| 9 | 253.0 | 253.0 | 252.8 | 252.9 | 253.0 | 253.0 | 253.1 | 253.0 | 252.9 |
| 10 | 253.4 | 253.6 | 253.2 | 253.4 | 253.4 | 253.3 | 253.5 | 253.4 | 253.6 |
Worked example 4: attribute agreement, braze fillet appearance
Constructed data, not a real production run.
Three inspectors each make an accept (0) / reject (1) call on the same 30 braze-fillet parts, twice each, in randomized order and without seeing their own or each other's prior calls. Six of the 30 parts were deliberately chosen to be borderline; a known "standard" call exists for every part (established separately, e.g. by a senior inspector or a destructive check), which is what makes it possible to score effectiveness and not just inter-inspector agreement.
| Part | Standard | A-1 | A-2 | B-1 | B-2 | C-1 | C-2 |
|---|---|---|---|---|---|---|---|
| 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 3 | 0 | 0 | 1 | 0 | 0 | 0 | 0 |
| 4 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| 5 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| 6 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 7 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| 8 | 0 | 1 | 0 | 0 | 0 | 0 | 1 |
| 9 | 0 | 1 | 1 | 1 | 1 | 1 | 1 |
| 10 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 11 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 12 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| 13 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 14 | 0 | 0 | 0 | 0 | 1 | 1 | 0 |
| 15 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 16 | 0 | 1 | 0 | 0 | 1 | 0 | 1 |
| 17 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| 18 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 19 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| 20 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 21 | 0 | 0 | 1 | 0 | 0 | 0 | 0 |
| 22 | 1 | 1 | 1 | 1 | 1 | 1 | 0 |
| 23 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 24 | 0 | 1 | 0 | 0 | 0 | 0 | 0 |
| 25 | 1 | 1 | 0 | 1 | 1 | 1 | 0 |
| 26 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
| 27 | 0 | 0 | 1 | 0 | 0 | 1 | 0 |
| 28 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| 29 | 0 | 1 | 1 | 0 | 1 | 0 | 0 |
| 30 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| Inspector | Within-inspector repeatability (both trials agree) | Effectiveness vs. standard | Cohen's κ vs. standard |
|---|---|---|---|
| A | 0.767 | 0.817 | 0.645 |
| B | 0.833 | 0.883 | 0.772 |
| C | 0.767 | 0.850 | 0.772 |
Fleiss' κ among the three inspectors (trial 1 only) = 0.626 By the Landis and Koch benchmarks, 0.626 is "substantial" agreement among the inspectors, and each individual inspector's kappa against the known standard (0.645 to 0.772) is at or above that same band. The inspectors agree with the truth slightly more than they agree with each other, because their mistakes on the six borderline parts do not always land on the same side.
None of the three numbers here is alarming on its own, but an 85 % overall hit rate on a visual accept/reject call, with kappa in the "substantial" rather than "almost perfect" band, is exactly the signal that a written, photo-referenced operational definition of "reject" (Module 5) would likely move the needle more than more training alone.
Common mistakes
- Running a capability study before checking the gauge. Consequence: an apparent process problem that is really a measurement problem, and weeks of the wrong corrective action (the illustrative failure above). Fix: MSA is the first Measure-phase deliverable, not an afterthought.
- Treating repeatability and reproducibility as interchangeable. Consequence: a reproducibility problem (training, definition) gets "fixed" by buying new equipment, or a repeatability problem (a worn fixture) gets "fixed" by retraining operators who were never the issue. Fix: read the ANOVA table; it tells you which one is large.
- Checking bias at only one point in the range. Consequence: a gauge that is accurate at the calibration point but biased at the extremes passes a single-point check and then mis-measures every part near the tolerance limits, which is exactly where the capability decision matters most. Fix: a linearity study across the full operating range, not a single bias check.
- Dropping the part × operator interaction from the ANOVA without checking its significance. Consequence: a real operator-by-part effect (one appraiser reads certain part geometries differently than others do) gets silently folded into repeatability, understating reproducibility. Fix: test the interaction first; pool it only when its p-value says it is not distinguishable from noise.
- Quoting %Study Variation and %Tolerance as if they always agree. Consequence: a gauge can look "acceptable" by one figure and "conditional" by the other, particularly when the parts sampled for the study do not span the full tolerance range. Fix: report both, and understand that %Tolerance needs parts spanning the range to mean anything.
- Treating the AIAG 10/30 bands and ndc ≥ 5 as universal laws. Consequence: an over-strict rejection of a gauge that is genuinely adequate for a non-critical characteristic, or an under-strict acceptance of one used on a safety characteristic. Fix: these are guidelines; state them as such and let the characteristic's risk set the bar.
- Running an attribute study with no known standard. Consequence: inter-rater kappa can be computed, but "how often is the team actually right" cannot be, which is usually the question that matters most. Fix: establish a reference/expert call for every part in the study before the appraisers rate them.
- Reading raw percent agreement instead of kappa. Consequence: two appraisers who both call almost everything "accept" look like they agree 90 %+ of the time even if their judgment adds nothing, because chance agreement alone is high when one category dominates. Fix: always report kappa, which corrects for exactly this.
Exercises
Exercise 1: bias study on a second gauge
Constructed data, not a real production run.
A dial caliper used to check a keyway width is checked against a certified reference standard of 6.015 mm, 12 repeat readings in one sitting.
| Reading | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mm | 6.010 | 6.011 | 6.010 | 6.005 | 6.012 | 6.010 | 6.007 | 6.011 | 6.010 | 6.010 | 6.009 | 6.011 |
Tasks. (a) Compute the mean reading and the bias against the 6.015 mm reference. (b) Is the bias distinguishable from zero? (c) The keyway's tolerance is 6.000 +0.030/0 mm (a 0.030 mm tolerance band). Express the bias as a percentage of tolerance and say whether it is a practical concern.
Show the worked solution
(a) Mean = 6.0097 mm. Bias = 6.0097 − 6.015 = −0.0053 mm (the gauge reads low).
(b) One-sample t-test against 6.015: t = -9.61 on 11 degrees of freedom, p < 0.001. Yes — this is a real, repeatable negative bias, not chance scatter.
(c) 0.0053 / 0.030 = 17.7 % of tolerance. Unlike Worked example 1's bore-gauge bias (5.6 % of tolerance), this is large enough to matter: a part that is genuinely 0.005 mm inside the lower boundary of the keyway tolerance could be measured as passing when the true dimension is out of spec at the edge. This gauge needs recalibration before it is used for acceptance decisions.
Exercise 2: a smaller attribute agreement study
Constructed data, not a real production run.
Two inspectors call accept (0) / reject (1) on 20 connector-housing parts, twice each.
| Part | Standard | A-1 | A-2 | B-1 | B-2 |
|---|---|---|---|---|---|
| 1 | 0 | 0 | 0 | 0 | 0 |
| 2 | 1 | 1 | 1 | 1 | 1 |
| 3 | 0 | 0 | 0 | 0 | 0 |
| 4 | 0 | 0 | 0 | 0 | 0 |
| 5 | 0 | 0 | 0 | 0 | 0 |
| 6 | 1 | 1 | 0 | 1 | 1 |
| 7 | 1 | 1 | 1 | 1 | 1 |
| 8 | 0 | 0 | 0 | 0 | 0 |
| 9 | 0 | 0 | 0 | 1 | 0 |
| 10 | 0 | 0 | 0 | 1 | 0 |
| 11 | 0 | 0 | 0 | 0 | 0 |
| 12 | 1 | 1 | 1 | 1 | 1 |
| 13 | 0 | 0 | 0 | 0 | 0 |
| 14 | 0 | 0 | 0 | 0 | 0 |
| 15 | 0 | 0 | 0 | 0 | 0 |
| 16 | 1 | 1 | 0 | 1 | 1 |
| 17 | 0 | 0 | 0 | 0 | 0 |
| 18 | 0 | 0 | 0 | 1 | 0 |
| 19 | 0 | 0 | 0 | 0 | 0 |
| 20 | 0 | 1 | 1 | 0 | 1 |
Tasks. (a) Compute each inspector's within-inspector repeatability (fraction of parts where their two trials agree). (b) Compute the overall raw agreement with the standard, across both inspectors and both trials. (c) Compute Fleiss' kappa between the two inspectors (using trial 1 only) and classify it on the Landis and Koch scale.
Show the worked solution
(a) Inspector A: 0.90 (18 of 20 parts). Inspector B: 0.80 (16 of 20 parts) — B is visibly less repeatable than A, disagreeing with themself on parts 9, 10, 18 and 20.
(b) Overall effectiveness against the standard, all 80 calls: 0.90.
(c) Fleiss' κ (trial 1, both inspectors) = 0.560. On the Landis and Koch scale that falls in the "moderate" band (0.41 to 0.60), a step below Worked example 4's three-inspector study (0.626, "substantial"). With only two inspectors and 20 parts, this study is also smaller than a production attribute study would normally be run at; the point of this exercise is the calculation, not a claim that 20 parts is an adequate sample size for a real acceptance decision.
Quiz
Ten questions. Score 70 % or more to mark the module complete on this device.
Answer key
- c. Observed = process + measurement variation.
- b. Same part, same appraiser, same gauge.
- d. Linearity.
- a. Pool into repeatability or keep as reproducibility.
- c. Under 10 % / 10-30 % / over 30 %.
- b. Distinguishable groups the gauge can separate.
- d. Subtracts out chance agreement.
- a. Three or more raters.
- c. Training and operational definition.
- b. Diagnose the specific failed component and fix that.
Key takeaways
- Observed variation is process variation plus measurement variation; a capability index means nothing until the measurement share is known to be small.
- Bias, linearity and stability studies each ask a different question: is the gauge accurate on average, does that accuracy hold across the range, and does it hold over time.
- Repeatability (same appraiser, same gauge) and reproducibility (different appraisers) are different sources of variation with different fixes: instrument and fixturing for repeatability, training and operational definitions for reproducibility.
- A crossed gauge R&R ANOVA separates part, operator, interaction and repeatability variance; %Contribution, %Study Variation, %Tolerance and ndc each answer a slightly different question and can disagree.
- AIAG's %GRR bands (under 10 % / 10-30 % / over 30 %) and ndc ≥ 5 are guidelines subject to customer agreement, not statistical laws.
- Attribute agreement analysis uses Cohen's kappa (two raters) or Fleiss' kappa (three or more) to correct raw percent agreement for chance; without a known standard, effectiveness against the truth cannot be computed, only inter-rater agreement.
- A failed gauge study is a diagnosis, not a dead end: which number failed (bias, linearity, stability, repeatability, reproducibility) decides which fix applies.
References
All web sources accessed 2026-09-09 or 2026-09-10 as noted. Sources marked "secondary" were not read in the original by the course author; the claim is taken from the source shown.
- NIST/SEMATECH. e-Handbook of Statistical Methods, 2.4 "Gauge R&R studies" and 2.4.4 "Analysis of variability". NIST. https://www.itl.nist.gov/div898/handbook/mpc/section4/mpc4.htm
- Burdick, R. K., Borror, C. M., & Montgomery, D. C. (2003). A review of methods for measurement systems capability analysis. Journal of Quality Technology, 35(4), 342–354. https://www.tandfonline.com/doi/abs/10.1080/00224065.2003.11980232
- NIST/SEMATECH. e-Handbook of Statistical Methods, 2.4 "Gauge R&R studies" (nested and crossed variance-component estimation). NIST. https://www.itl.nist.gov/div898/handbook/mpc/section4/mpc4.htm
- AIAG. Measurement Systems Analysis (MSA) Reference Manual, 4th ed., June 2010. AIAG. https://www.aiag.org/training-and-resources/manuals/details/MSA-4 (the %GRR acceptance bands and ndc ≥ 5 guideline are confirmed via secondary sources [8], [9])
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://journals.sagepub.com/doi/10.1177/001316446002000104
- Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382. https://doi.org/10.1037/h0031619
- Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://pubmed.ncbi.nlm.nih.gov/843571/ (secondary: the benchmark scale quoted here is confirmed via citation-listing sources, not a direct reading of the paywalled original)
- SPC for Excel. "Acceptance criteria for measurement systems analysis." https://www.spcforexcel.com/knowledge/measurement-systems-analysis-gage-rr/acceptance-criteria-for-msa/ (secondary: restates the AIAG MSA manual's bands)
- QualityEngineer.ai. "Gauge R&R acceptance criteria: %GRR, NDC, and what AIAG MSA requires." https://app.qualityengineer.ai/blog/gauge-rr-acceptance-criteria (secondary: corroborates [8])