Module 9 · Analyze
Graphical analysis
Analyze opens with pictures, not tests, and that order is deliberate. A histogram, a Pareto chart, a scatter plot or a multi-vari chart can show a pattern in thirty seconds that a table of summary statistics hides completely, and looking first keeps you from running a sophisticated test on data that a two-minute plot would have told you not to trust. This module covers the graphical toolkit Analyze reaches for before Module 11's hypothesis tests: Pareto charts, box plots, scatter plots, stratification, multi-vari charts, and time series plots, with Anscombe's quartet as the standing argument for why the look-first discipline exists at all.
Learning objectives
- Build a Pareto chart and use it to separate the vital few contributors from the trivial many.
- Read a box plot's five-number summary and use it to compare groups at a glance.
- Use a scatter plot to check a relationship visually before any regression (Module 12) is run.
- Stratify a dataset by a suspected source of variation before drawing a conclusion from the pooled data.
- Build and interpret a multi-vari chart, separating positional, cyclical and temporal sources of variation.
- Explain, using Anscombe's quartet, why graphical analysis has to come before, not after, a statistical test.
Why this matters
Joseph Juran's 1975 note is a small piece of intellectual honesty worth remembering every time someone reaches for the "80/20 rule": Juran named the pattern of a few contributors accounting for most of a problem "the vital few and the trivial many," and attributed it to the economist Vilfredo Pareto's observations on wealth distribution. He later admitted the attribution was his own error — Pareto never made the general statement Juran's tool is named after — and published a "mea culpa" saying so in print.[1] The tool itself is unaffected: sorting causes by frequency and looking at where the cumulative line bends is exactly as useful as it always was. What the story is worth remembering for is the habit it models — check the thing you are about to repeat as received wisdom, even the name of your own tool — which is the same habit graphical analysis asks of the data itself, look before repeating a summary statistic as if it were the whole story.
Francis Anscombe made the sharper version of that argument in 1973, with four small, real datasets, each eleven points, each sharing the same mean, variance, correlation, and regression line to two or three decimal places — and each looking completely different on a scatter plot: one a clean linear relationship, one a clear curve that a straight-line fit misses entirely, one a perfect line broken by a single outlier, one a single vertical cluster with all its apparent "relationship" created by one extreme point.[2] Module 1 introduced this dataset; this module is where its lesson becomes a standing habit rather than a one-time demonstration.
Pareto charts
A Pareto chart sorts categories by frequency (or cost, or any other measure of impact) from largest to smallest, as bars, with a cumulative-percentage line overlaid. The chart answers one question directly: how many categories account for most of the problem? Ishikawa listed the Pareto chart among the seven basic quality tools precisely because that question is usually answerable at a glance, without any further analysis.[6]
Box plots
John Tukey's box plot draws a five-number summary — minimum, first quartile, median, third quartile, maximum — as a box (the interquartile range, Q1 to Q3) with a line at the median and whiskers extending to the most extreme points within a set distance of the box, conventionally 1.5 times the interquartile range; points beyond that are drawn individually as candidate outliers.[3] A box plot's real strength is comparison: several groups' box plots side by side show shifted medians, unequal spreads, and skew, all at once, in a way a table of means and standard deviations does not communicate nearly as quickly.
Scatter plots
A scatter plot of one variable against another is the single most direct check of whether a linear relationship (Module 12) is even a reasonable model to fit: a plot showing a curve, a broken relationship, or a pattern driven by one or two points warns against trusting a correlation coefficient or a regression line computed blind. Anscombe's quartet is the standing proof that the summary numbers alone cannot tell you this — only the picture can.[5]
Stratification
Stratification means splitting a dataset by a suspected source of variation — shift, machine, operator, lot, cavity — before drawing a conclusion from the pooled data, rather than after. Module 5's rational subgrouping lesson is stratification applied to control charts specifically; the same discipline applies just as directly to a histogram, a box plot, or a simple average: a pooled histogram that looks unremarkable can be two cleanly separated, non-overlapping distributions once split by the right factor, exactly as Worked example 3 below shows for bore diameter by shift.
Multi-vari charts
A multi-vari chart is a structured, deliberately stratified sampling plan built to answer one question before any formal test is run: of the variation this process shows, how much comes from within one part (positional), how much from part to part at roughly the same time (cyclical), and how much from time to time (temporal)? The method originates with Leonard Seder's 1950 two-part article.[4] The sampling plan is deliberate: measure several positions on each of several parts, repeated at several points across a shift or a day, so that all three sources of variation are represented in one dataset instead of confounded together.
Time series plots
A simple run chart — the measured value plotted in time order, nothing more — is the plot every other graphical tool in this module implicitly depends on being meaningful: it is what confirms (or refutes) that the time order in the data is real and worth stratifying by, and it is the direct precursor to the control chart Modules 16 and 17 build from the same time-ordered data with limits added.
When to stop looking and start testing
Every tool in this module generates hypotheses. None of them tests one. A box plot that shows shift 3 sitting higher than shifts 1 and 2 is a reason to ask what shift 3 does differently; it is not evidence that the difference is larger than the scatter you would see if all three shifts were identical. That question has an answer, and Module 11 computes it. The order matters in both directions: testing without looking gets you a p-value on a relationship whose shape you never checked, and looking without testing gets you a confident story built on a difference that four more parts would have erased.
The practical rule is that a picture earns a test when a decision depends on it. If the box plot sends you to look at shift 3's setup sheet and you find the fixture is different, the picture has done its whole job and no test is needed, because the physical difference is now the evidence. If the picture is going to be the evidence, in a report, in a supplier discussion, or in a decision to spend money, it needs a test and an interval beside it.
Reading a plot honestly
Because these plots are so persuasive, the ways they mislead are worth naming. Three come up constantly in real reports:
- A truncated vertical axis. Starting the axis at 11.99 instead of zero is often the right choice for a measured characteristic, because zero is meaningless and the interesting variation is in the fourth decimal. It is the wrong choice for a bar chart of counts, where the bar's length is the whole point and truncating it multiplies an apparent difference. Read the axis before the bars, every time.
- Two vertical axes on one plot. Any two series can be made to look correlated by choosing the two scales, and any two can be made to look unrelated the same way. If the relationship between two variables is the point, plot them against each other in a scatter plot, where the reader can see the relationship rather than the scaling.
- Unequal bins and dropped points. A histogram with one wide bin at the tail hides the tail; a scatter plot with the "obviously wrong" point removed hides whatever produced it. Outliers are investigated and annotated with a documented cause, never quietly deleted, and a plot that has had points removed says so in its caption.
The same discipline applies to what you plot in the first place. A Pareto chart of defect categories is only as good as the categories, which usually come from an operator's drop-down list rather than from any analysis; a large "other" category or one enormous vague category is a data collection problem (Module 5) surfacing as a chart, not a finding.
Worked examples
Worked example 1: Pareto of leak-test reject causes
The data below is a constructed example, not a real production run.
169 leak-test rejects on the brazed heat-exchanger line over one quarter, root cause assigned at teardown.
| Cause | Count | % | Cumulative % |
|---|---|---|---|
| Header joint | 87 | 51.5 | 51.5 |
| Fin-to-tube joint | 34 | 20.1 | 71.6 |
| End cap seal | 21 | 12.4 | 84.0 |
| Braze void | 12 | 7.1 | 91.1 |
| Flux residue | 9 | 5.3 | 96.4 |
| Handling damage | 6 | 3.6 | 100.0 |
The top 3 of 6 categories — header joint, fin-to-tube joint, and end cap seal — account for 84.0 % of all rejects. That is not the folklore "20 % of categories, 80 % of the problem" ratio exactly (3 of 6 is 50 % of the categories), and it does not need to be: the header joint alone, one cause out of six, is already worth more corrective-action attention than the other five combined.
Worked example 2: multi-vari chart, braze fillet width
Constructed data, not a real production run.
Braze fillet width at the header joint, measured at 3 positions around each fillet, on 4 consecutive parts sampled at each of 5 points spanning one shift.
| Part | Position | T1 | T2 | T3 | T4 | T5 |
|---|---|---|---|---|---|---|
| P1 | 1 | 1.169 | 1.216 | 1.226 | 1.213 | 1.245 |
| P1 | 2 | 1.173 | 1.219 | 1.217 | 1.226 | 1.237 |
| P1 | 3 | 1.165 | 1.218 | 1.210 | 1.218 | 1.241 |
| P2 | 1 | 1.178 | 1.214 | 1.208 | 1.239 | 1.249 |
| P2 | 2 | 1.185 | 1.212 | 1.200 | 1.232 | 1.252 |
| P2 | 3 | 1.185 | 1.212 | 1.197 | 1.228 | 1.256 |
| P3 | 1 | 1.194 | 1.222 | 1.203 | 1.222 | 1.239 |
| P3 | 2 | 1.191 | 1.214 | 1.203 | 1.229 | 1.233 |
| P3 | 3 | 1.194 | 1.228 | 1.206 | 1.224 | 1.251 |
| P4 | 1 | 1.189 | 1.208 | 1.213 | 1.234 | 1.246 |
| P4 | 2 | 1.194 | 1.217 | 1.216 | 1.234 | 1.239 |
| P4 | 3 | 1.191 | 1.223 | 1.212 | 1.234 | 1.248 |
Reading down any one column (within a part, across its 3 positions) shows only a little scatter. Reading across a row's near neighbours (part to part, same time point) shows somewhat more. Reading across the table left to right (time point to time point) shows a clear, steady climb from about 1.18 mm at T1 to about 1.24 mm at T5 — consistent with a furnace running warmer as the shift goes on.
Cyclical R̄ = 0.0150 mm, σ̂cyclical = R̄/d₂(4) = 0.0073 mm
Temporal σ̂ = sample sd of the 5 time-period means = 0.0225 mm Time-period means: T1 1.1840, T2 1.2169, T3 1.2092, T4 1.2278, T5 1.2447 mm.
Temporal variation accounts for about two-thirds of the total, roughly three times either of the other two sources. The multi-vari chart does not, by itself, prove the furnace is the cause — it points at "something that changes over the course of a shift" as the dominant source and tells the team exactly where the next round of investigation (a furnace temperature log, Module 16's control chart applied to that log) should look first, instead of starting with fixture-to-fixture or part-to-part theories that this chart already shows are comparatively minor.
Worked example 3: box plot by shift
Constructed data, not a real production run.
Bore diameter, Ø12.000 ± 0.020 mm, 30 consecutive parts from each of 3 shifts.
| Shift | Min | Q1 | Median | Q3 | Max |
|---|---|---|---|---|---|
| 1 | 11.986 | 11.9947 | 11.9995 | 12.005 | 12.009 |
| 2 | 11.986 | 11.994 | 12.0005 | 12.00625 | 12.016 |
| 3 | 11.987 | 12.002 | 12.0085 | 12.01625 | 12.028 |
Pooled into one histogram, these 90 values would show a somewhat wide, slightly right-skewed distribution and nothing more alarming than that. Stratified by shift, shift 3's box sits visibly higher (median 12.008 mm against shift 1's 12.000 mm) and taller (s = 0.0091 mm against shift 1's 0.0064 mm), and its single highest reading, 12.028 mm, is over the 12.020 mm upper specification limit — a genuine out-of-spec part that the pooled view would have buried in the tail of one wide-looking distribution instead of flagging as "whichever shift this came from is worth a visit."
Show the 90 bore diameters behind Table 3 and Figure 3
| Part | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1–10 | 11.99 | 11.992 | 11.992 | 11.998 | 11.986 | 11.999 | 11.994 | 12.005 | 12.006 | 12.008 |
| 11–20 | 12.005 | 12.0 | 12.005 | 12.009 | 11.996 | 12.004 | 12.0 | 12.009 | 11.995 | 11.998 |
| 21–30 | 12.002 | 12.002 | 11.99 | 12.002 | 11.999 | 11.999 | 11.999 | 12.001 | 11.989 | 12.009 |
| Part | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1–10 | 11.996 | 11.988 | 12.001 | 12.01 | 11.998 | 12.01 | 12.01 | 11.996 | 11.986 | 11.994 |
| 11–20 | 12.008 | 11.996 | 11.992 | 11.993 | 12.001 | 12.001 | 12.006 | 12.007 | 11.995 | 12.0 |
| 21–30 | 11.994 | 12.001 | 11.994 | 11.994 | 12.014 | 12.001 | 12.005 | 12.016 | 12.006 | 12.0 |
| Part | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| 1–10 | 12.006 | 11.999 | 12.016 | 12.003 | 12.022 | 12.003 | 12.003 | 12.002 | 12.012 | 12.008 |
| 11–20 | 12.018 | 12.001 | 12.009 | 12.023 | 12.004 | 12.011 | 12.02 | 12.011 | 12.0 | 11.987 |
| 21–30 | 12.028 | 12.01 | 12.017 | 12.002 | 12.022 | 12.011 | 12.006 | 11.998 | 12.009 | 12.002 |
Common mistakes
- Running a hypothesis test before looking at a plot of the data. Consequence: Anscombe's quartet — identical summary statistics, wildly different underlying relationships; a test run blind can confidently answer the wrong question. Fix: plot first, always, no exceptions for data that "looks routine."
- Treating the 80/20 ratio as a rule the Pareto chart must obey. Consequence: real data that shows, say, three categories needed to reach 80 % gets treated as if the tool failed, when the tool did exactly its job. Fix: read the cumulative line for what it actually shows, not for a folklore ratio.
- Comparing groups by their means alone, skipping the box plot. Consequence: two groups with nearly identical means but very different spreads, or different shapes, look "the same" in a table and completely different once plotted. Fix: box plot before, or alongside, any table of summary statistics.
- Fitting a regression line to a scatter plot that was never looked at. Consequence: Anscombe's set 2 (a curve) and set 3 (one outlier driving the whole fit) both produce a "good" linear fit by the numbers while the plot shows the model is wrong. Fix: look at the scatter plot before trusting r or r² (Module 12).
- Skipping stratification because the pooled histogram "looks fine." Consequence: Worked example 3 — a pooled distribution can look unremarkable while hiding two or three cleanly separated sub-populations. Fix: stratify by every suspected source (shift, machine, operator, lot) before concluding a distribution is homogeneous.
- Running a multi-vari study without a deliberate sampling plan. Consequence: sampling parts and positions haphazardly confounds the three sources of variation instead of separating them, and the chart cannot say which one dominates. Fix: decide the number of positions, parts and time points in advance, matching Module 5's sampling-plan discipline.
- Treating a multi-vari chart's result as a finished root cause, not a pointer. Consequence: "temporal variation dominates" identifies where to look next (what changes over time in this process), not the specific mechanism; stopping at the chart skips the investigation it was supposed to start. Fix: use the dominant source to scope the next round of data collection or a designed experiment (Module 13), not as the final answer.
Exercises
Exercise 1: a second Pareto chart
Constructed example, not a real production run.
A wire-harness sub-assembly logged 112 defects over one month at final inspection: missing clip 58, mislabelled harness 24, crimp pull-out 16, wrong connector 9, damaged insulation 5. Tasks. (a) Compute each category's percentage and cumulative percentage. (b) How many categories are needed to reach at least 80 % of the total? (c) Name one risk of stopping the investigation after fixing only the single largest category.
Show the worked solution
| Cause | Count | % | Cumulative % |
|---|---|---|---|
| Missing clip | 58 | 51.8 | 51.8 |
| Mislabelled harness | 24 | 21.4 | 73.2 |
| Crimp pull-out | 16 | 14.3 | 87.5 |
| Wrong connector | 9 | 8.0 | 95.5 |
| Damaged insulation | 5 | 4.5 | 100.0 |
(b) 3 categories (missing clip, mislabelled harness, crimp pull-out) reach 87.5 %, already past 80 %.
(c) Fixing only "missing clip" leaves 46.2 % of defects (the remaining four categories) completely unaddressed; the vital-few logic identifies where to start, not where to stop.
Exercise 2: a multi-vari chart with a different dominant source
Constructed example, not a real production run.
Clamp fixture flatness deviation (mm), 3 positions per part, 3 parts per time point, 3 time points across a shift. Tasks. (a) Without seeing the numbers, name one plausible physical cause each for positional, cyclical, and temporal variation dominating in a fixture-flatness study. (b) The computed shares are positional 62.2 %, cyclical 31.9 %, temporal 5.8 %. Which source dominates, and what would you investigate first?
Show the worked solution
(a) Positional: the fixture clamps unevenly, so different spots on the same part read differently every time. Cyclical: parts themselves vary (incoming material or a prior process step). Temporal: something drifts over the shift (temperature, fixture wear, operator fatigue).
(b) Positional dominates, at nearly two-thirds of the total (0.0198 mm of σ̂, against 0.0102 mm cyclical and only 0.0018 mm temporal). This is the opposite pattern from Worked example 2 (where temporal dominated): the fixture itself, not the furnace or the incoming parts, is where the next investigation should start — consistent with an unevenly clamping fixture, one of the (a) hypotheses.
Show the fixture flatness data (3 time points × 3 parts × 3 positions)
| Time point | Part | Position 1 | Position 2 | Position 3 |
|---|---|---|---|---|
| 1 | P1 | 0.791 | 0.755 | 0.831 |
| 1 | P2 | 0.799 | 0.819 | 0.81 |
| 1 | P3 | 0.816 | 0.793 | 0.792 |
| 2 | P1 | 0.798 | 0.809 | 0.789 |
| 2 | P2 | 0.782 | 0.813 | 0.802 |
| 2 | P3 | 0.806 | 0.781 | 0.813 |
| 3 | P1 | 0.768 | 0.773 | 0.815 |
| 3 | P2 | 0.826 | 0.82 | 0.81 |
| 3 | P3 | 0.804 | 0.823 | 0.787 |
Quiz
Ten questions. Score 70 % or more to mark the module complete on this device.
Answer key
- c. Same statistics, different structures, visible only by plotting.
- b. Juran's own later correction.
- d. How many categories reach a given cumulative share.
- a. Q1 to Q3.
- c. Shift 3 higher, more variable, one part over USL.
- b. Positional, cyclical, temporal.
- d. Points at where to look next.
- a. Range-based decomposition, R̄/d₂ style.
- c. Splitting by a suspected source before conclusions.
- b. Time-ordered plot, precursor to a control chart.
Key takeaways
- Graphical analysis comes before hypothesis testing (Module 11) because summary statistics alone can look identical for data that is structurally very different — Anscombe's quartet is the standing proof.
- A Pareto chart sorts categories by impact and reads the cumulative line; there is no rule that exactly 80 % must come from 20 % of categories, and even the tool's own name is a documented misattribution Juran corrected himself.
- A box plot's five-number summary compares groups' medians, spreads and skew at a glance, faster and more completely than a table of means and standard deviations.
- Stratification — splitting data by a suspected source before drawing a conclusion — can reveal that a pooled, unremarkable-looking distribution is really two or three separated sub-populations.
- A multi-vari chart uses a deliberate sampling plan to separate positional, cyclical and temporal variation, pointing at where to look next rather than delivering a finished root cause.
- A time series (run) plot is the precursor to a control chart: it is what confirms that time order in the data is real and worth stratifying by.
References
All web sources accessed 2026-09-09 unless noted.
- Juran, J. M. (1975). The non-Pareto principle; mea culpa. Quality Progress, 8(5), 8–9. https://www.juran.com/wp-content/uploads/2021/03/The-Non-Pareto-Principle-1974.pdf
- Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27(1), 17–21. Values as reproduced in Wikipedia and the R
datasetspackage documentation (Module 1, S-E35, S-E36). https://www.sjsu.edu/faculty/gerstman/StatPrimer/anscombe1973.pdf - Tukey, J. W. (1977). Exploratory Data Analysis. Addison-Wesley. https://www.scirp.org/reference/referencespapers?referenceid=1482121 (catalogue-level; box plot origin)
- Seder, L. A. (1950). Diagnosis with diagrams, Parts I and II. Industrial Quality Control, 6(4) and 6(5). https://asq.org/quality-progress/articles/case-studies/the-multi-vari-chart-an-underutilized-quality-tool?id=d23574b44b3d49b88e9dcdddc4785cbb (secondary: the original 1950 articles were not located; the attribution and method description are via this ASQ case study and corroborating sources)
- NIST/SEMATECH. e-Handbook of Statistical Methods, 4.4.4 "How can I tell if a model fits my data?". NIST. https://www.itl.nist.gov/div898/handbook/pmd/section4/pmd44.htm
- Ishikawa, K. (1976). Guide to Quality Control. Asian Productivity Organization. https://openlibrary.org/books/OL4595409M/Guide_to_quality_control (catalogue-level; the seven basic quality tools, of which the Pareto chart is one)
- ASQ. Certified Six Sigma Green Belt (CSSGB) Body of Knowledge Map 2014–2022. ASQ, 2022. https://www.asq.org/cert/resource/pdf/certification/2022-CSSGB-BoK-Map.pdf (the Green Belt scope this module is written against; no single claim rests on it)