Module 5 · Measure
Data collection and sampling
A gauge that passes its MSA study (Module 4) can still produce a useless dataset, because how the data is collected matters as much as how accurately each individual reading is taken. Two engineers can run gauge R&R on the identical instrument, then go collect process data in two different ways, and end up with two different control charts, two different capability verdicts, and two different conclusions about the same stable process. The difference is not the gauge. It is the sampling plan and, above everything else in this module, how the data gets split into subgroups. This module covers operational definitions, sampling plans, rational subgrouping, data collection sheets, and the everyday data quality failures that undermine a study before Analyze ever begins.
Learning objectives
- Write an operational definition precise enough that two people applying it independently reach the same result.
- Distinguish a sampling plan that represents the process from one that is merely convenient to collect.
- Explain rational subgrouping and why the subgrouping choice decides what a control chart can and cannot detect.
- Demonstrate, from data, how subgrouping the same measurements two different ways produces two different control limits and two different verdicts.
- Design a data collection sheet that captures the context a later analysis will need.
- Recognise the everyday data quality failures — coarse rounding, mixed streams, sorting before measuring — that quietly distort a study.
Why this matters
Donald Wheeler's account of rational subgrouping makes a claim stronger than most engineers expect: the subgrouping decision is not a detail to settle after the data is collected, it is the decision that determines what question a control chart is capable of answering at all.[2] A subgroup is supposed to represent "some small region of space, or time, or product" in which the process was doing only what it routinely does; the within-subgroup variation that estimates a chart's sigma should be nothing but that routine, common-cause noise. If a subgroup accidentally spans two different sources of variation — two cavities of a mold, two machines, two shifts, a tool change — the chart's own estimate of "routine" variation gets contaminated by variation that was never routine, and the chart's limits widen to absorb it. A real difference between the two sources then hides inside artificially wide limits instead of showing up as the signal it actually is.
The scenario below is a constructed illustration, not a real incident.
No fully documented public case study for this specific module was found in the sources searched for this course, so the illustration below is a labelled constructed scenario rather than a real one. A machining cell samples five consecutive parts every hour for its X̄-R chart. Two nominally identical lathes feed the same conveyor, and the sampler just grabs "the next five parts." Because the two machines run continuously and interleave on the conveyor, roughly half of every subgroup comes from each lathe. The chart has run stable for months. What it cannot show, by construction, is that lathe 2 has quietly drifted 0.018 mm high of lathe 1: that difference is real, it is happening on every single part, and it is completely invisible on this chart, because it is baked into "normal" subgroup-to-subgroup noise rather than appearing as a between-subgroup signal. The chart is not lying. It is answering a question — "does the combined two-machine stream drift over time?" — that nobody meant to ask, instead of the question that mattered: "are these two machines producing the same thing?"
Operational definitions
An operational definition states, in terms specific enough to leave no room for judgement, exactly how a characteristic is measured and how the result is classified. "Reject if the fillet looks thin" is not an operational definition; "reject if the fillet width at the marked reference point measures below 1.2 mm on the gauge described in WI-114" is. Deming argued that most disputes over data are really disputes over definitions that were never made operational, and that a numerical specification is meaningless without one, because two people can apply the same nominal rule to the same part and reach opposite conclusions if the rule leaves any of the measurement method, the reference point, the environment, or the classification boundary unstated.[1], [7] Module 4's attribute agreement study is a direct, quantitative test of whether an operational definition is doing its job: a low kappa is usually evidence that the definition, not the inspectors, needs work.
Sampling plans
A sampling plan states what will be measured, how often, by whom, and from where in the process, decided before data collection starts rather than improvised by whoever happens to be free. Convenience sampling — measuring whatever is easiest to reach, whenever there is a spare moment — is the most common failure, because it silently trades representativeness for ease: parts from the end of a shift, parts near the inspection station, parts a particular operator happens to run. A sampling plan should say explicitly which population the sample is meant to represent (every part? every cavity? every shift?) and confirm the sampling method actually reaches all of it. The ASQ Green Belt Body of Knowledge places sampling plans and data collection squarely inside Measure, alongside rational subgrouping, as the foundation the rest of the phase depends on.[4]
Rational subgrouping and why it decides everything downstream
The formal rule, in Wheeler's phrasing: choose subgroups so that the variation within a subgroup is only the routine variation you want the chart's limits to represent, and let any variation you want the chart to be able to detect show up between subgroups.[2], [5] This is a design decision, not an afterthought, and it has to be made before data collection starts, because it cannot be fixed by more sophisticated analysis afterward — the information that a naive subgrouping throws away by blending two sources together is simply gone.
Between-subgroup variation is what the X̄ chart is built to detect. Consecutive-in-time subgroups (the default choice) are rational when the process runs as one stream. When a process has more than one stream — cavities, spindles, operators, machines — consecutive-in-time subgroups are usually not rational, because they mix the streams inside every subgroup.
The standard formulas for the chart (Module 16 builds them in full) do not change: Ā and R̄ from the subgroups, limits at A₂R̄ either side of the grand mean.[3] What changes entirely is what those limits mean, depending on which values ended up in the same subgroup.
Data collection sheets
A data collection sheet exists to capture the context a later analysis will need, not just the measurement itself. At minimum: the value, the time it was taken, who took it, which machine/cavity/spindle/operator produced the part, and the gauge used. Skipping the context columns is the single most common reason a promising-looking dataset turns out to be unusable three weeks later, when someone asks "was this before or after the tool change?" and nobody recorded it.
Show the 80 micrometer readings (0.001 mm resolution)
| Part | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1–10 | 6.503 | 6.507 | 6.503 | 6.49 | 6.507 | 6.504 | 6.496 | 6.505 | 6.503 | 6.502 |
| 11–20 | 6.5 | 6.504 | 6.494 | 6.499 | 6.496 | 6.505 | 6.5 | 6.498 | 6.494 | 6.498 |
| 21–30 | 6.5 | 6.498 | 6.51 | 6.508 | 6.478 | 6.485 | 6.499 | 6.497 | 6.502 | 6.502 |
| 31–40 | 6.517 | 6.491 | 6.497 | 6.516 | 6.505 | 6.505 | 6.496 | 6.487 | 6.501 | 6.501 |
| 41–50 | 6.49 | 6.495 | 6.499 | 6.492 | 6.499 | 6.501 | 6.5 | 6.496 | 6.505 | 6.507 |
| 51–60 | 6.503 | 6.493 | 6.506 | 6.496 | 6.507 | 6.491 | 6.507 | 6.5 | 6.49 | 6.497 |
| 61–70 | 6.5 | 6.502 | 6.492 | 6.491 | 6.502 | 6.496 | 6.502 | 6.506 | 6.487 | 6.502 |
| 71–80 | 6.51 | 6.498 | 6.494 | 6.506 | 6.502 | 6.507 | 6.497 | 6.488 | 6.499 | 6.496 |
Common data quality failures
Three failures recur often enough to name individually, all of them things a data collection sheet and a sampling plan should have prevented, and none of them fixable by better statistics after the fact.
Rounding to the tolerance
Recording a measurement more coarsely than the process's own spread requires throws away real variation and replaces it with quantisation noise. A common rule of thumb: gauge resolution should divide the tolerance into at least ten increments (echoing the same "gauge before process" logic as Module 4); recording to fewer than that starts to matter, and recording to only two or three distinguishable values across the whole spread of the data makes a control chart's range-based sigma estimate unreliable, because a moving range or subgroup range computed from a handful of repeated, rounded values is not really measuring the same thing as one computed from continuous data.
Mixed streams
Combining two sources — cavities, machines, operators, shifts, incoming lots — into a single stream before subgrouping, sampling, or even charting is the same failure as poor rational subgrouping, but it is not limited to control charts: a capability study, a hypothesis test, or a simple histogram can all be quietly wrong if the data is secretly two different populations stacked together. The fix is always the same: know your sources before you pool your data, and check whether pooling is throwing away information you actually need.
Sorting before measuring
Parts arrive at an inspection station in a tote, get sorted by size or appearance before measurement "for convenience," and the original production time order is gone. Rational subgrouping and control charting both depend on time order (or some other rational basis) being preserved; once it is destroyed, no subgroup can honestly represent "a small region of time," and any control chart built afterward is really just a histogram in disguise, incapable of detecting a time-based shift because the time axis no longer means anything.
Worked examples
Worked example 1: the same 100 measurements, subgrouped two ways
The data below is a constructed example, not a real production run.
A two-cavity injection mould produces a bracket hole, Ø8.000 ± 0.020 mm, bore gauge to 0.001 mm. Cavity A runs centred on nominal; cavity B has always run about 0.010 mm high, a known, small, stable difference nobody has bothered to chase down because both cavities individually make parts well inside the tolerance. Parts leave the mould strictly alternating cavity A, B, A, B, …, 100 parts in production order. The sampler takes the next five parts off the line every hour for the X̄-R chart — the naive, consecutive-in-time subgrouping below, which is rational only if there is one stream, and here there are two.
| Subgroup | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| 1 | 8.002 | 8.014 | 8.002 | 8.003 | 8.005 |
| 2 | 8.012 | 7.997 | 8.013 | 8.002 | 8.011 |
| 3 | 8.000 | 8.013 | 7.996 | 8.009 | 7.998 |
| 4 | 8.013 | 8.000 | 8.009 | 7.996 | 8.009 |
| 5 | 8.000 | 8.009 | 8.006 | 8.015 | 7.986 |
| 6 | 8.001 | 7.999 | 8.008 | 8.001 | 8.011 |
| 7 | 8.011 | 8.004 | 7.998 | 8.020 | 8.003 |
| 8 | 8.013 | 7.997 | 8.002 | 8.001 | 8.011 |
| 9 | 7.994 | 8.007 | 8.000 | 8.005 | 8.000 |
| 10 | 8.010 | 8.000 | 8.007 | 8.003 | 8.014 |
| 11 | 8.002 | 8.006 | 8.004 | 8.007 | 8.004 |
| 12 | 8.005 | 8.005 | 8.010 | 7.994 | 8.008 |
| 13 | 8.000 | 8.011 | 7.995 | 8.004 | 8.001 |
| 14 | 8.008 | 8.001 | 8.014 | 7.992 | 8.011 |
| 15 | 8.006 | 8.009 | 7.996 | 8.014 | 8.001 |
| 16 | 8.014 | 7.998 | 8.003 | 7.999 | 8.008 |
| 17 | 8.004 | 8.011 | 7.992 | 8.004 | 8.004 |
| 18 | 8.013 | 7.997 | 8.010 | 8.002 | 8.012 |
| 19 | 8.004 | 8.011 | 8.000 | 8.009 | 8.005 |
| 20 | 7.999 | 7.999 | 8.010 | 7.993 | 8.012 |
UCL(X̄) = 8.0140 mm, LCL(X̄) = 7.9952 mm Stability verdict: no point beyond the limits, no run rule fires. The chart says this process is stable and has been for the whole study.
Now the same 100 numbers, not one of them changed, regrouped so that every subgroup is a single cavity: subgroups 1–10 are cavity A's 50 parts in their own production order, subgroups 11–20 are cavity B's.
| Subgroup | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| 1 (A) | 8.002 | 8.002 | 8.005 | 7.997 | 8.002 |
| 2 (A) | 8.000 | 7.996 | 7.998 | 8.000 | 7.996 |
| 3 (A) | 8.000 | 8.006 | 7.986 | 7.999 | 8.001 |
| 4 (A) | 8.011 | 7.998 | 8.003 | 7.997 | 8.001 |
| 5 (A) | 7.994 | 8.000 | 8.000 | 8.000 | 8.003 |
| 6 (A) | 8.002 | 8.004 | 8.004 | 8.005 | 7.994 |
| 7 (A) | 8.000 | 7.995 | 8.001 | 8.001 | 7.992 |
| 8 (A) | 8.006 | 7.996 | 8.001 | 7.998 | 7.999 |
| 9 (A) | 8.004 | 7.992 | 8.004 | 7.997 | 8.002 |
| 10 (A) | 8.004 | 8.000 | 8.005 | 7.999 | 7.993 |
| 11 (B) | 8.014 | 8.003 | 8.012 | 8.013 | 8.011 |
| 12 (B) | 8.013 | 8.009 | 8.013 | 8.009 | 8.009 |
| 13 (B) | 8.009 | 8.015 | 8.001 | 8.008 | 8.011 |
| 14 (B) | 8.004 | 8.020 | 8.013 | 8.002 | 8.011 |
| 15 (B) | 8.007 | 8.005 | 8.010 | 8.007 | 8.014 |
| 16 (B) | 8.006 | 8.007 | 8.005 | 8.010 | 8.008 |
| 17 (B) | 8.011 | 8.004 | 8.008 | 8.014 | 8.011 |
| 18 (B) | 8.009 | 8.014 | 8.014 | 8.003 | 8.008 |
| 19 (B) | 8.011 | 8.004 | 8.013 | 8.010 | 8.012 |
| 20 (B) | 8.011 | 8.009 | 7.999 | 8.010 | 8.012 |
UCL(X̄) = 8.0108 mm, LCL(X̄) = 7.9985 mm Same grand mean (both groupings average the identical 100 numbers), but R̄ drops by about a third, because it is no longer measuring "cavity-to-cavity difference" mixed in with true part-to-part noise — it is measuring only the true part-to-part noise, cavity by cavity, which is what a range chart is supposed to estimate.
Correctly subgrouped, the chart is not stable. The primary 3-sigma rule catches only 3 of the 20 points (subgroups 2, 3 and 7, all cavity A, dipping just past the now-tighter LCL by chance), but the supplementary run rules leave no doubt: eight subgroups in a row on the low side (subgroups 3 through 10) and eight in a row on the high side (13 through 20) both trip the "8 in a row on one side of the centre line" rule, and the two-of-three-beyond-2-sigma and four-of-five-beyond-1-sigma rules fire repeatedly through both halves of the chart. Anyone looking at Figure-1-style output from this grouping would see the step immediately; anyone looking only at Table 1's naive chart would see nothing at all. The real, persistent 0.010 mm cavity-to-cavity difference was never a mystery hiding in the process. It was hiding in the subgrouping.
Control chart builder — naive subgrouping
Pre-loaded with Table 1 (production order, mixed cavities).
Control chart builder — rational subgrouping
Pre-loaded with Table 2 (the same 100 values, one cavity per subgroup). Try turning on the Nelson rules.
Worked example 2: rounding loss
Constructed data, not a real production run.
A locating pin, Ø6.500 ± 0.030 mm, is measured 80 times. The true readings, taken with a digital micrometer and recorded to 0.001 mm, are compared with the same 80 parts read from a dial caliper and recorded only to 0.01 mm — ten times coarser than the micrometer, and coarse enough, relative to this process's own tight spread, to matter.
| Resolution | Mean (mm) | s (mm) | Distinct values observed |
|---|---|---|---|
| Fine, 0.001 mm (micrometer) | 6.4994 | 0.00677 | most of the 80 unique |
| Coarse, 0.01 mm (dial caliper) | 6.4998 | 0.00763 | 5 |
The mean barely moves. The standard deviation does not: it inflates from 0.00677 to 0.00763 mm, about 13 % higher, and the entire 80-part sample now lands on only five distinguishable values. Rounding to a resolution w adds independent, roughly uniformly distributed quantisation error with variance w²/12, so the observed variance is approximately the true variance plus w²/12 — here (0.01)²/12 ≈ 0.0000083, enough on its own to move a process this tight by a visible amount. The distortion runs in only one direction: coarse rounding can only add apparent variation, never remove it, and it also throws away the fine structure a control chart or a capability study needs to tell a genuinely tight process from a merely adequate one.
Descriptive statistics and histogram
Pre-loaded with the coarse (0.01 mm) readings from Table 3. Compare against the fine readings by pasting them in.
Worked example 3: a sampling plan for a three-shift line
Constructed example, drawn for teaching.
The capstone's leak-test operation (Module 19) runs three shifts a day, seven days a week. A sampling plan for its ongoing p chart has to state, in advance, every one of the items below — leaving any of them to be decided ad hoc on the floor is how a chart quietly stops representing what it claims to represent.
| What is sampled | Every unit that completes the leak test (100 % inspection at this station; the "sample" is the full population, so no selection bias is possible here). |
|---|---|
| Subgroup boundary | One subgroup per shift (three per day), not per hour and not per day — shift boundaries are rational here because shift-to-shift differences (operator, furnace warm-up state) are exactly the kind of variation the chart should be able to detect between subgroups. |
| Subgroup size | Variable, equal to that shift's actual production count (a p chart, not an np chart, because n is not constant). |
| Who records it | The leak-test operator, logged automatically by the test station at end of shift; no manual transcription step. |
| What is NOT in scope | Units reworked and re-tested are logged under the shift they were re-tested in, not the shift they were originally built in, and flagged as rework so Analyze can separate first-pass results from rework results on request. |
| Review cadence | Control limits recalculated monthly from the trailing 20 shifts; a special cause investigated the same shift it appears, not batched for the monthly review. |
Notice what makes this a plan and not just a description of what already happens: every decision (shift-based subgrouping, where rework goes, how often limits get recalculated) was made deliberately, in writing, before data collection, specifically so that no single sampler's convenience or a supervisor's after-the-fact judgement call quietly changes what the chart is measuring from one week to the next.
Common mistakes
- Subgrouping by consecutive production order without asking whether there is more than one stream. Consequence: Worked example 1 — a real, persistent difference hides inside artificially wide limits and the chart reports false stability. Fix: identify every stream (cavity, spindle, machine, operator) before deciding how to subgroup, not after.
- Recording data more coarsely than the process's own spread requires. Consequence: Worked example 2 — inflated apparent variation, a handful of distinguishable values standing in for what should be a continuous distribution, and range-based sigma estimates that no longer mean what they are supposed to mean. Fix: resolution should divide the tolerance (or the expected process spread, whichever is tighter) into at least ten steps.
- An operational definition that leaves the measurement method unstated. Consequence: two people, or the same person on two days, get different answers from nominally the same rule, and nobody can tell whether a chart signal is a process change or a definition drifting. Fix: write down the method, the reference point and the classification boundary, not just the pass/fail criterion.
- Sorting or batching parts before measuring them. Consequence: production time order is destroyed, and with it any chart's ability to detect a time-based shift; the "control chart" becomes a histogram wearing a time axis it no longer earns. Fix: measure in the order parts were made, or record the true order alongside the measurement if physical sorting cannot be avoided.
- Convenience sampling dressed up as a sampling plan. Consequence: whatever is easiest to reach (end of shift, near the inspection station, one operator's output) is systematically over-represented, and the sample silently stops representing the population it claims to. Fix: state which population the plan represents and confirm the method actually reaches all of it.
- Data collection sheets with no context columns. Consequence: a promising dataset becomes unusable weeks later because nobody recorded which machine, which operator, or what changed and when. Fix: capture context at the time of measurement; it cannot be reconstructed afterward.
- Assuming rational subgrouping is a one-time decision. Consequence: a process that gains a second stream (a new cavity added, a second shift started) keeps using the old subgrouping scheme, and the new source's variation gets silently absorbed the same way cavity B was in Worked example 1. Fix: revisit the subgrouping plan whenever the process configuration changes.
Exercises
Exercise 1: two machines, naive vs rational
Constructed example, not a real production run.
A shaft shoulder diameter, Ø25.000 ± 0.040 mm, is machined on two nominally identical lathes that feed the same conveyor in strict alternation, 40 parts total. Naive subgrouping takes 8 consecutive subgroups of 5 in production order (mixing both lathes in every subgroup); rational subgrouping regroups the same 40 values so subgroups 1–4 are lathe 1 and subgroups 5–8 are lathe 2.
Tasks. (a) Without computing anything, predict which grouping will show the smaller R̄. (b) Compute R̄ for both groupings and check your prediction. (c) Which grouping's X̄ chart is more likely to be unstable, and why is that the useful outcome rather than a problem to explain away?
Show the worked solution
(a) Rational, because within-lathe variation alone should be smaller than variation that mixes two lathes running at different levels.
(b) Naive R̄ = 0.0328 mm. Rational R̄ = 0.0196 mm — about 40 % smaller, confirming the prediction.
(c) The rational grouping's chart is the one worth trusting: it is not stable, while the naive grouping reports a falsely reassuring "stable, no signal." The rational grouping correctly flags that something worth investigating (the two lathes running at different levels) is present. An unstable chart from a rationally subgrouped study is not a failure of the study; it is the study doing its job.
Show the 40 bore diameters, in production order
| Subgroup | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| 1 | 24.992 | 25.005 | 24.998 | 25.022 | 25.011 |
| 2 | 25.019 | 24.994 | 25.01 | 25.007 | 25.034 |
| 3 | 25.003 | 25.006 | 24.99 | 25.034 | 25.002 |
| 4 | 25.001 | 24.999 | 25.006 | 24.994 | 25.013 |
| 5 | 24.993 | 25.024 | 24.999 | 25.012 | 25.004 |
| 6 | 25.026 | 24.984 | 25.015 | 24.99 | 25.016 |
| 7 | 24.987 | 25.018 | 25.0 | 25.015 | 24.99 |
| 8 | 25.014 | 24.989 | 25.004 | 25.002 | 25.007 |
Exercise 2: spot the data collection problems
Constructed scenario, not a real project.
A new hire is asked to collect 100 readings of a fill weight for a capability study. Their plan: "I'll grab whatever containers are sitting on the finished-goods pallet when I have a free ten minutes each day, sort them by how full they look so the light ones are easy to spot, and write down the weight to the nearest gram since that's what the customer spec is in." Task. Identify every data quality failure in this plan and say what it will do to the resulting study.
Show the worked solution
- Convenience sampling: "whatever is sitting on the pallet" when free is not tied to production time order or any defined population; it will over-represent whatever accumulates near the pallet at that time of day.
- Sorting before measuring: sorting by apparent fullness destroys production order entirely; any time-based shift becomes undetectable, and a "control chart" built from this data cannot function as one.
- Rounding to the tolerance, not the process: recording to the nearest gram because that is the spec's unit, rather than to a resolution that resolves the process's actual spread, risks exactly the sd inflation and value-clumping shown in Worked example 2 if the true spread is only a few grams.
The fix for all three: a written sampling plan (which containers, on what schedule, in what order) agreed before data collection starts, measurement in production order with order preserved even if physical handling is out of order, and a resolution chosen from the process's own expected spread, not from the specification.
Quiz
Ten questions. Score 70 % or more to mark the module complete on this device.
Answer key
- c. Two people, same result.
- b. Within-subgroup = routine only; between-subgroup = what the chart should detect.
- a. R̄ decreased; the real difference was revealed.
- d. Wider limits, hiding a real difference.
- c. Inflates the estimated sd.
- b. Destroys the time order the chart needs.
- a. Over-represents whatever is easiest to reach.
- d. Context: time, who, which source, which gauge.
- c. Decided in advance.
- b. Reviewed; the old scheme may now mix streams.
Key takeaways
- An operational definition is precise enough that two people applying it independently reach the same result; Module 4's attribute agreement study is a direct test of whether one is working.
- A sampling plan states what, how often, by whom and from where, decided in advance; convenience sampling silently trades representativeness for ease.
- Rational subgrouping's rule: within-subgroup variation should be only routine, common-cause variation; anything you want the chart to be able to detect should show up between subgroups.
- The same 100 measurements, subgrouped two different ways, can produce two different R̄s and two opposite stability verdicts; the subgrouping choice, not more sophisticated analysis afterward, decides what a chart can detect.
- Recording data more coarsely than the process's own spread requires inflates the apparent standard deviation and destroys the fine structure a chart or capability study needs.
- Sorting or batching parts before measuring destroys production time order, and with it, a control chart's ability to detect a time-based shift.
- A data collection sheet should capture context (time, source, gauge, operator), not just the value; context not captured at the time cannot be reconstructed later.
References
All web sources accessed 2026-09-09 unless noted. Sources marked "secondary" were not read in the original by the course author; the claim is taken from the source shown.
- Deming, W. E. (1975). On probability as a basis for action. The American Statistician, 29(4), 146–152. https://deming.org/wp-content/uploads/2020/06/On-Probability-As-a-Basis-For-Action-1975.pdf
- Wheeler, D. J. "Rational subgrouping: the conceptual foundation of process behavior charts." Quality Digest, 1 June 2015 (manuscript 282); "Rational sampling," 1 July 2015 (manuscript 283). https://spcpress.com/pdf/DJW282.pdf
- NIST/SEMATECH. e-Handbook of Statistical Methods, 6.3.2.1 "Shewhart X-bar and R and S control charts". NIST. https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc321.htm
- ASQ. Certified Six Sigma Green Belt (CSSGB) Body of Knowledge Map 2014–2022. ASQ, 2022. https://www.asq.org/cert/resource/pdf/certification/2022-CSSGB-BoK-Map.pdf
- Wheeler, D. J., & Chambers, D. S. (2010). Understanding Statistical Process Control, 3rd ed. SPC Press. https://www.spcpress.com/book_understanding_statistical_process_control.php (catalogue-level; general reference for rational subgrouping's role in SPC)
- Montgomery, D. C. (2019). Introduction to Statistical Quality Control, 8th ed. Wiley. https://www.wiley.com/en-us/Introduction+to+Statistical+Quality+Control,+8th+Edition-p-9781119399308 (secondary: catalogue-level; further reading, no claim on this page rests on it)
- Deming, W. E. (1986). Out of the Crisis. MIT Center for Advanced Engineering Study. https://archive.org/details/outofcrisisquali00demi (secondary: catalogue-level; operational definitions are discussed at length in this book, not independently verified by page this session)