Module 4 · Measure

Measurement systems analysis

Before Measure collects a single production value, it has to answer a question that has nothing to do with the process: can the gauge be trusted? Every number in this course, from a mean to a Cpk to a p-value, is a number the gauge reported, not a number the process actually produced. If the gauge's own variation is a large share of what gets measured, no amount of statistical sophistication downstream can tell the difference between a process problem and a measurement problem. This module covers the full measurement systems analysis (MSA) toolkit: bias, linearity and stability studies, repeatability and reproducibility, crossed gauge R&R by the ANOVA method, attribute agreement analysis, and what to actually do when a gauge fails.

Learning objectives

Why this matters

NIST's e-Handbook frames the problem in one line: a measurement is not the true value of a characteristic, it is the true value plus measurement error, and that error has its own repeatability, reproducibility and stability that must be characterized before the measurement can be trusted for anything else.[1] Burdick, Borror and Montgomery's 2003 review of gauge capability methods makes the same point from the ANOVA side: the variance an inspector observes in a set of measurements is not the part-to-part variance alone, it is part variance plus operator variance plus a part-by-operator interaction plus pure measurement error, all superimposed, and a two-factor ANOVA is the standard way to pull them apart.[2] Total observed variance decomposes as σ²observed = σ²process + σ²measurement. A process with genuinely tight, capable production can still produce a terrible-looking Cpk if the gauge measuring it is noisy; a genuinely loose process can look artificially capable if the gauge is too coarse to resolve the scatter. Module 3's crankshaft case is a reminder of the same discipline from the other direction: a DMAIC project that reports a Cp or a Cpk owes the reader confidence that the number describes the process and not the gauge.

The scenario below is a constructed illustration, not a real incident.

A common, entirely avoidable failure: a machining cell gets chartered for a capability problem because its bore diameter Cpk has been drifting for three months. Two weeks into Measure, someone finally runs a gauge R&R on the bore gauge that has been reporting every one of those numbers, and finds %GRR of 45 % of the tolerance — the gauge itself, not the machining process, explains most of the apparent variation. The three months of capability trend data, every control chart built from it, and the machine adjustments made in response to "special causes" that were actually gauge noise, all have to be set aside. The fix (a worn gauge anvil, replaced for the cost of a spare part) takes an afternoon. The MSA study that should have caught it in week one, run last, cost the project most of a quarter. Running the gauge study first is not bureaucracy; it is the cheapest possible insurance against exactly this outcome.

Why the gauge comes before the process

Every characteristic this course measures — a bore diameter, a pull strength, a leak-test result — is observed through some instrument: a micrometer, a load cell, a go/no-go fixture, a trained eye. The observed value is never the part's true dimension; it is the true dimension plus whatever the gauge and the person using it add or subtract on that particular reading. If that added variation is small relative to the process's own spread, it can be ignored. If it is not small, every downstream number is compromised: a capability index computed from noisy data understates true capability, a control chart built on a noisy gauge has wider limits than the process itself would produce and so misses real shifts, and a hypothesis test comparing two conditions loses power because gauge noise is added to both sides. MSA answers one question before any of that work starts: how much of what I am about to measure is the part, and how much is the measuring?

Bias, linearity, and stability

These three studies ask whether the gauge, on average, reads the true value (bias), whether that accuracy holds across the whole range the gauge is used over (linearity), and whether it stays that way over time (stability). All three need one thing the process itself does not require: a certified reference standard with a known true value.

Bias

Bias = mean of repeated readings on the reference standard − the reference standard's known value A one-sample t-test against zero (Module 11 covers the test in full; here it answers one question: is this bias larger than chance would produce from a gauge with no true bias?).

A small bias is not automatically a problem; a bias that is both statistically significant and large relative to the tolerance is. AIAG's guidance, echoed across the MSA literature, is to report bias as a percentage of the process tolerance so an engineer can judge practical significance alongside the t-test's statistical significance.

Linearity

A gauge can be unbiased on average across its whole range and still be biased low at one end and high at the other; averaging those two errors together can hide a real problem. A linearity study measures several certified reference standards spread across the gauge's operating range, computes the bias at each level, and regresses bias against the reference value. A flat line at bias = 0 is the ideal; a non-zero, statistically significant slope means the gauge's error itself depends on where in the range it is reading.

Bias(reference value) = intercept + slope × reference value The slope is the quantity of interest: it says how many millimetres of bias change per millimetre of reference value. A slope indistinguishable from zero (by the regression's own t-test on the slope, Module 12) means linearity is not a concern over the range tested.

Stability

Bias and linearity are snapshots. Stability asks whether that snapshot holds up over weeks of use: does the same reference standard, measured once a shift, drift as the gauge wears, gets recalibrated, or is handled by different people over time? An I-MR chart (Module 16 builds the general tool; here it is applied to gauge readings instead of process output) of repeated reference readings answers this the same way a process control chart answers it for a process: points beyond the 3-sigma limits, or a run, mean something changed.

Repeatability and reproducibility

Repeatability is the variation in repeated measurements of the same part, by the same appraiser, with the same gauge, close together in time — the gauge's own inherent noise, sometimes called equipment variation (EV). Reproducibility is the additional variation that shows up when different appraisers measure the same parts with the same gauge — operator-to-operator variation, sometimes called appraiser variation (AV). A gauge can be highly repeatable (an individual operator gets nearly identical readings every time) and still poorly reproducible (three different operators, each internally consistent, systematically disagree with each other), and a study has to check both, because they call for different fixes: a repeatability problem is usually the instrument or the fixture, a reproducibility problem is usually training, ergonomics, or an ambiguous operational definition (Module 5 covers operational definitions directly).

Crossed gauge R&R by the ANOVA method

A crossed study has every appraiser measure every part the same number of times, in randomized order, usually blind to their own prior readings. "Crossed" means every appraiser-part combination is observed, which is what lets a two-way ANOVA separate part variation, operator variation, and the part-by-operator interaction (does a particular operator read particular parts differently than other operators do) from pure repeatability error, following Burdick, Borror and Montgomery's review of the method.[2]

Repeatability (equipment variation) = the ANOVA error term's mean square (or, if the part×operator interaction is not significant, pooled with it)
Reproducibility (appraiser variation) = the operator variance component + the interaction variance component
Total gauge R&R variance = repeatability variance + reproducibility variance Variance components are recovered from the ANOVA mean squares by subtraction (each mean square's expected value under the random-effects model minus the next one down, divided by the appropriate replication count), the same logic NIST uses for nested gauge studies.[3] A first F-test on the interaction term decides whether to keep it or pool it into repeatability; AIAG and Minitab both default to pooling when that test's p-value exceeds 0.05.
%Contribution = 100 × (variance component / total variance)
%Study Variation = 100 × (standard deviation / total standard deviation), using k = 6 (99.73 % of a normal distribution) or k = 5.15 (99 %) times the standard deviation as "study variation"
%Tolerance = 100 × (k × standard deviation / tolerance)
ndc (number of distinct categories) = floor(1.41 × σpart / σGRR) %Contribution uses variances (so the gauge and part shares sum to 100); %Study Variation and %Tolerance use standard deviations, so they do not sum to 100 the same way and can tell different stories about the same gauge, particularly when the part sample does not span the full tolerance. ndc estimates how many truly distinguishable groups the gauge can separate parts into — AIAG's rule of thumb is 5 or more.

AIAG's Measurement Systems Analysis reference manual sets out the acceptance bands used across the industry: %GRR (the larger of %Study Variation and %Tolerance) under 10 % is generally acceptable, 10 to 30 % may be acceptable depending on the application and the cost of a better gauge, and over 30 % is unacceptable.[4][8], [9] These bands, and the ndc ≥ 5 rule of thumb, are guidelines a customer can tighten or relax by agreement, not laws of statistics; a safety characteristic may reasonably demand better than 10 %, and a cosmetic one may reasonably tolerate more.

Attribute agreement analysis

Not every characteristic is measured on a continuous scale. A visual accept/reject call, a go/no-go gauge, a pass/fail leak test — all of these need their own kind of measurement system study, because "repeatability" for a discrete judgment is not a standard deviation, it is whether the same appraiser gives the same answer twice, and whether different appraisers agree with each other and, if a known correct answer exists, with the truth.

Cohen's kappa measures chance-corrected agreement between two raters on a categorical judgment.[5] Fleiss' kappa extends the idea to three or more raters.[6] Both start from the same idea: raw percent agreement overstates how good a rater is, because two raters who both call 90 % of parts "accept" will agree with each other about 82 % of the time by chance alone, even if neither one is looking at the part. Kappa subtracts out that chance agreement.

κ = (po − pe) / (1 − pe) po is the observed proportion of agreement; pe is the proportion of agreement expected if both raters were assigning categories at random, at their own observed marginal rates. κ = 1 is perfect agreement, κ = 0 is exactly what chance would produce, and κ < 0 is worse than chance.

Landis and Koch's widely used benchmark scale reads κ ≤ 0 as poor, 0.01 to 0.20 as slight, 0.21 to 0.40 as fair, 0.41 to 0.60 as moderate, 0.61 to 0.80 as substantial, and 0.81 to 1.00 as almost perfect agreement.[7] Like the AIAG %GRR bands, this is a rule of thumb the literature has converged on, not a statistical law; a safety inspection may need "substantial" or better before the study is accepted at all.

What to do with a bad gauge

A failed study is a diagnosis, not a dead end, and which fix applies depends on which number failed.

What never fixes a bad measurement system: widening the acceptance criteria until the gauge passes, or quietly reporting the process capability anyway with a footnote. If the gauge cannot be trusted, no number computed from it can be trusted either, and that has to be the conclusion stated to the sponsor, however unwelcome.

Worked examples

Worked example 1: bias and linearity, the bore gauge

The data below is a constructed example, not a real production run.

The bore gauge used throughout Modules 7 and 16 for the Ø12.000 ± 0.025 mm running example gets its own MSA study here. First, bias: a certified reference standard with a known diameter of 12.010 mm is measured 15 times in one sitting.

Table 1. Bias study, 15 readings of a 12.010 mm reference standard (constructed data).
Reading123456789101112131415
mm12.01212.01112.01212.01312.01512.01312.01212.01212.01412.01512.01312.01112.01112.01512.013
Mean reading = 12.0128 mm    Bias = 12.0128 − 12.010 = 0.0028 mm A one-sample t-test against the reference value (Module 11) asks whether this is distinguishable from zero: t = 7.61 on 14 degrees of freedom, p < 0.001. The bias is small in absolute terms but far too consistent across 15 readings to be chance; it is a real, repeatable offset.

0.0028 mm against a 0.050 mm tolerance is 5.6 % of tolerance — worth a calibration adjustment, not an emergency. Next, linearity: five certified reference standards spanning the tolerance, four repeat readings each.

Table 2. Linearity study, 5 reference levels x 4 replicates (constructed data). Bias = measured − reference.
Reference (mm)Measured (mm)Bias (mm)
11.98511.984−0.001
11.98511.9850.000
11.98511.984−0.001
11.98511.982−0.003
11.99511.9960.001
11.99511.9960.001
11.99511.994−0.001
11.99511.9960.001
12.00512.0070.002
12.00512.0070.002
12.00512.0070.002
12.00512.0070.002
12.01512.0170.002
12.01512.0180.003
12.01512.0170.002
12.01512.0190.004
12.02512.0290.004
12.02512.0290.004
12.02512.0280.003
12.02512.0290.004

Average bias runs from about −0.00125 mm at the low end (11.985 mm) to about 0.00375 mm at the high end (12.025 mm) — small numbers, but they move in one direction as the reference value increases, which is exactly what a linearity problem looks like.

Bias = intercept + slope × reference value Slope = 0.1225 (SE 0.0134, t = 9.14, p < 0.001), R² = 0.823. The slope is small — about 0.12 µm of extra bias per millimetre of reference value — but with only 0.0015 mm of repeatability noise per reading, four replicates at each of five levels is enough to detect it clearly.

Both findings point the same way: this gauge reads a little high, increasingly so toward the top of the range. A single-point recalibration at mid-range would fix the average bias but leave the linearity trend untouched; a multi-point calibration correction, or gauge maintenance across the full range, is the appropriate fix (see "What to do with a bad gauge" above).

Worked example 2: stability, the same reference standard over time

Constructed data, not a real production run.

The same 12.010 mm reference standard is measured once at the start of each of 24 shifts, to see whether the bias found above holds steady or drifts.

Table 3. Stability study, one reading per shift for 24 shifts (constructed data).
Shift123456789101112131415161718192021222324
mm12.01312.01412.01312.01112.01412.01312.01212.01412.01312.01312.01312.01412.01212.01312.01212.01412.01312.01212.01212.01212.01312.01212.01512.014
x̄ = 12.0130 mm, MR̄ = 0.0012 mm, σ̂ = MR̄/1.128 = 0.0010 mm UCL = 12.0161 mm, LCL = 12.0098 mm, UCL(MR) = 0.0038 mm.

No point beyond the limits, no run, no trend: this gauge is stable over the 24 shifts studied, even though it carries the small bias and linearity effects found above. Stability and accuracy are different questions; a gauge can be perfectly consistent about being slightly wrong.

Worked example 3: crossed gauge R&R, two gauges compared

Constructed data, not a real production run.

Two crossed studies, same design (10 parts × 3 operators × 3 trials, randomized order), on two different gauges: the bore gauge above, and a checkweigher used for fill-weight inspection (spec 248.0 to 254.0 g, tolerance 6.0 g). Both ANOVA tables test the part × operator interaction first; in both cases the interaction is not significant (bore gauge p = 0.229; checkweigher p = 0.162) and is pooled into repeatability before the variance components are computed.

Table 4. Gauge R&R comparison, two gauges, both 10 x 3 x 3 crossed studies.
QuantityBore gauge (Ø12.000 ± 0.025 mm)Checkweigher (248.0-254.0 g)
%Contribution (GRR)5.7 %0.5 %
%Study Variation (GRR)23.9 %7.0 %
%Tolerance (GRR)29.8 %11.4 %
ndc520
AIAG verdictConditional (10-30 % band)Acceptable (under 10 %)

Same study design, same analysis method, two very different verdicts. The bore gauge's %GRR sits in AIAG's conditional band on both the study-variation and tolerance figures, with ndc right at the minimum of 5 — usable, but a candidate for the fixturing and resolution review described above, especially given the bias and linearity findings on this same gauge in Worked example 1. The checkweigher is comfortably acceptable on every figure. This is also why Module 7's bore capability numbers (Cwk about 1.06) carry a caveat this module makes explicit: with %GRR near 24 % of study variation, some of the "process" variation that capability index reports on is genuinely the gauge's.

Gauge R&R (crossed, ANOVA method)

Pre-loaded with the bore gauge study above (10 parts × 3 operators × 3 trials). Paste your own measurements to run a different study.

Show the checkweigher study data (10 parts × 3 operators × 3 trials, 90 readings)
Table A1. Checkweigher gauge R&R: fill weight (g), 10 parts, operators A, B and C, three trials each (constructed data). Column labels are operator and trial, so B2 is operator B's second reading of that part.
PartA1A2A3B1B2B3C1C2C3
1248.6248.6248.6248.7248.7248.6248.6248.6248.7
2249.2249.4249.4248.9249.0249.2249.2249.3249.3
3250.0249.6249.7250.0249.8249.8249.7249.5249.8
4250.0249.9249.9250.0249.9250.0250.0250.0250.0
5250.9251.0250.9250.8251.0250.8251.0250.8251.0
6251.3251.2251.3251.3251.4251.2251.2251.4251.3
7251.7251.8251.5251.8251.9251.7251.6251.8251.8
8252.5252.4252.2252.4252.4252.5252.5252.3252.3
9253.0253.0252.8252.9253.0253.0253.1253.0252.9
10253.4253.6253.2253.4253.4253.3253.5253.4253.6

Worked example 4: attribute agreement, braze fillet appearance

Constructed data, not a real production run.

Three inspectors each make an accept (0) / reject (1) call on the same 30 braze-fillet parts, twice each, in randomized order and without seeing their own or each other's prior calls. Six of the 30 parts were deliberately chosen to be borderline; a known "standard" call exists for every part (established separately, e.g. by a senior inspector or a destructive check), which is what makes it possible to score effectiveness and not just inter-inspector agreement.

Table 5. Attribute agreement study, 3 inspectors x 30 parts x 2 trials, 0 = accept, 1 = reject (constructed data).
PartStandardA-1A-2B-1B-2C-1C-2
10000000
20000000
30010000
40001000
51111111
60000000
71111111
80100001
90111111
100000000
110000000
121111111
130000000
140000110
150000000
160100101
171111111
180000000
191111111
200000000
210010000
221111110
230000000
240100000
251101110
261111111
270010010
280000001
290110100
300001000
Table 6. Attribute agreement results, by inspector.
InspectorWithin-inspector repeatability (both trials agree)Effectiveness vs. standardCohen's κ vs. standard
A0.7670.8170.645
B0.8330.8830.772
C0.7670.8500.772
Overall raw agreement with the standard (all 180 calls) = 0.85
Fleiss' κ among the three inspectors (trial 1 only) = 0.626 By the Landis and Koch benchmarks, 0.626 is "substantial" agreement among the inspectors, and each individual inspector's kappa against the known standard (0.645 to 0.772) is at or above that same band. The inspectors agree with the truth slightly more than they agree with each other, because their mistakes on the six borderline parts do not always land on the same side.

None of the three numbers here is alarming on its own, but an 85 % overall hit rate on a visual accept/reject call, with kappa in the "substantial" rather than "almost perfect" band, is exactly the signal that a written, photo-referenced operational definition of "reject" (Module 5) would likely move the needle more than more training alone.

Common mistakes

  1. Running a capability study before checking the gauge. Consequence: an apparent process problem that is really a measurement problem, and weeks of the wrong corrective action (the illustrative failure above). Fix: MSA is the first Measure-phase deliverable, not an afterthought.
  2. Treating repeatability and reproducibility as interchangeable. Consequence: a reproducibility problem (training, definition) gets "fixed" by buying new equipment, or a repeatability problem (a worn fixture) gets "fixed" by retraining operators who were never the issue. Fix: read the ANOVA table; it tells you which one is large.
  3. Checking bias at only one point in the range. Consequence: a gauge that is accurate at the calibration point but biased at the extremes passes a single-point check and then mis-measures every part near the tolerance limits, which is exactly where the capability decision matters most. Fix: a linearity study across the full operating range, not a single bias check.
  4. Dropping the part × operator interaction from the ANOVA without checking its significance. Consequence: a real operator-by-part effect (one appraiser reads certain part geometries differently than others do) gets silently folded into repeatability, understating reproducibility. Fix: test the interaction first; pool it only when its p-value says it is not distinguishable from noise.
  5. Quoting %Study Variation and %Tolerance as if they always agree. Consequence: a gauge can look "acceptable" by one figure and "conditional" by the other, particularly when the parts sampled for the study do not span the full tolerance range. Fix: report both, and understand that %Tolerance needs parts spanning the range to mean anything.
  6. Treating the AIAG 10/30 bands and ndc ≥ 5 as universal laws. Consequence: an over-strict rejection of a gauge that is genuinely adequate for a non-critical characteristic, or an under-strict acceptance of one used on a safety characteristic. Fix: these are guidelines; state them as such and let the characteristic's risk set the bar.
  7. Running an attribute study with no known standard. Consequence: inter-rater kappa can be computed, but "how often is the team actually right" cannot be, which is usually the question that matters most. Fix: establish a reference/expert call for every part in the study before the appraisers rate them.
  8. Reading raw percent agreement instead of kappa. Consequence: two appraisers who both call almost everything "accept" look like they agree 90 %+ of the time even if their judgment adds nothing, because chance agreement alone is high when one category dominates. Fix: always report kappa, which corrects for exactly this.

Exercises

Exercise 1: bias study on a second gauge

Constructed data, not a real production run.

A dial caliper used to check a keyway width is checked against a certified reference standard of 6.015 mm, 12 repeat readings in one sitting.

Exercise 1 data. 12 readings of a 6.015 mm reference standard (constructed data).
Reading123456789101112
mm6.0106.0116.0106.0056.0126.0106.0076.0116.0106.0106.0096.011

Tasks. (a) Compute the mean reading and the bias against the 6.015 mm reference. (b) Is the bias distinguishable from zero? (c) The keyway's tolerance is 6.000 +0.030/0 mm (a 0.030 mm tolerance band). Express the bias as a percentage of tolerance and say whether it is a practical concern.

Show the worked solution

(a) Mean = 6.0097 mm. Bias = 6.0097 − 6.015 = −0.0053 mm (the gauge reads low).

(b) One-sample t-test against 6.015: t = -9.61 on 11 degrees of freedom, p < 0.001. Yes — this is a real, repeatable negative bias, not chance scatter.

(c) 0.0053 / 0.030 = 17.7 % of tolerance. Unlike Worked example 1's bore-gauge bias (5.6 % of tolerance), this is large enough to matter: a part that is genuinely 0.005 mm inside the lower boundary of the keyway tolerance could be measured as passing when the true dimension is out of spec at the edge. This gauge needs recalibration before it is used for acceptance decisions.

Exercise 2: a smaller attribute agreement study

Constructed data, not a real production run.

Two inspectors call accept (0) / reject (1) on 20 connector-housing parts, twice each.

Exercise 2 data. 2 inspectors x 20 parts x 2 trials, 0 = accept, 1 = reject (constructed data).
PartStandardA-1A-2B-1B-2
100000
211111
300000
400000
500000
611011
711111
800000
900010
1000010
1100000
1211111
1300000
1400000
1500000
1611011
1700000
1800010
1900000
2001101

Tasks. (a) Compute each inspector's within-inspector repeatability (fraction of parts where their two trials agree). (b) Compute the overall raw agreement with the standard, across both inspectors and both trials. (c) Compute Fleiss' kappa between the two inspectors (using trial 1 only) and classify it on the Landis and Koch scale.

Show the worked solution

(a) Inspector A: 0.90 (18 of 20 parts). Inspector B: 0.80 (16 of 20 parts) — B is visibly less repeatable than A, disagreeing with themself on parts 9, 10, 18 and 20.

(b) Overall effectiveness against the standard, all 80 calls: 0.90.

(c) Fleiss' κ (trial 1, both inspectors) = 0.560. On the Landis and Koch scale that falls in the "moderate" band (0.41 to 0.60), a step below Worked example 4's three-inspector study (0.626, "substantial"). With only two inspectors and 20 parts, this study is also smaller than a production attribute study would normally be run at; the point of this exercise is the calculation, not a claim that 20 parts is an adequate sample size for a real acceptance decision.

Quiz

Ten questions. Score 70 % or more to mark the module complete on this device.

1. A measurement system analysis is run before a capability study because
2. Repeatability is best described as
3. A gauge is unbiased on average across its range but reads low at the low end and high at the high end. This is a
4. In a crossed gauge R&R ANOVA, the part × operator interaction term is tested first in order to
5. AIAG's %GRR acceptance guideline is generally stated as
6. The number of distinct categories, ndc, estimates
7. Cohen's kappa improves on raw percent agreement between two raters because it
8. Fleiss' kappa, compared with Cohen's kappa, is used when
9. A gauge R&R study shows very high reproducibility variation (large operator effect) but low repeatability variation. The most appropriate first response is to
10. If a gauge study fails its acceptance criteria, the correct response is to
Answer key
  1. c. Observed = process + measurement variation.
  2. b. Same part, same appraiser, same gauge.
  3. d. Linearity.
  4. a. Pool into repeatability or keep as reproducibility.
  5. c. Under 10 % / 10-30 % / over 30 %.
  6. b. Distinguishable groups the gauge can separate.
  7. d. Subtracts out chance agreement.
  8. a. Three or more raters.
  9. c. Training and operational definition.
  10. b. Diagnose the specific failed component and fix that.

Key takeaways

References

All web sources accessed 2026-09-09 or 2026-09-10 as noted. Sources marked "secondary" were not read in the original by the course author; the claim is taken from the source shown.

  1. NIST/SEMATECH. e-Handbook of Statistical Methods, 2.4 "Gauge R&R studies" and 2.4.4 "Analysis of variability". NIST. https://www.itl.nist.gov/div898/handbook/mpc/section4/mpc4.htm
  2. Burdick, R. K., Borror, C. M., & Montgomery, D. C. (2003). A review of methods for measurement systems capability analysis. Journal of Quality Technology, 35(4), 342–354. https://www.tandfonline.com/doi/abs/10.1080/00224065.2003.11980232
  3. NIST/SEMATECH. e-Handbook of Statistical Methods, 2.4 "Gauge R&R studies" (nested and crossed variance-component estimation). NIST. https://www.itl.nist.gov/div898/handbook/mpc/section4/mpc4.htm
  4. AIAG. Measurement Systems Analysis (MSA) Reference Manual, 4th ed., June 2010. AIAG. https://www.aiag.org/training-and-resources/manuals/details/MSA-4 (the %GRR acceptance bands and ndc ≥ 5 guideline are confirmed via secondary sources [8], [9])
  5. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://journals.sagepub.com/doi/10.1177/001316446002000104
  6. Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378–382. https://doi.org/10.1037/h0031619
  7. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://pubmed.ncbi.nlm.nih.gov/843571/ (secondary: the benchmark scale quoted here is confirmed via citation-listing sources, not a direct reading of the paywalled original)
  8. SPC for Excel. "Acceptance criteria for measurement systems analysis." https://www.spcforexcel.com/knowledge/measurement-systems-analysis-gage-rr/acceptance-criteria-for-msa/ (secondary: restates the AIAG MSA manual's bands)
  9. QualityEngineer.ai. "Gauge R&R acceptance criteria: %GRR, NDC, and what AIAG MSA requires." https://app.qualityengineer.ai/blog/gauge-rr-acceptance-criteria (secondary: corroborates [8])