Module 14 · Improve

Lean improvement tools

Analyze (Modules 9 through 12) finds and confirms a cause. Improve turns that finding into a change that actually sticks. This module covers the standard Lean toolkit for doing that on a shop floor: 5S and standard work as the stable foundation everything else depends on, poka-yoke for mistake-proofing, SMED for cutting changeover time, kaizen events for running a focused improvement effort, and a Pugh matrix for choosing among competing concepts. The last section is the one that matters most: how to pilot a change and prove it worked with a hypothesis test and a control chart, not a before-and-after bar chart.

Learning objectives

Why this matters

In 2007, BusinessWeek reported on what happened when 3M applied Six Sigma discipline broadly across the company under CEO James McNerney, including to research: thousands of scientists trained as Black Belts, standardized stage-gate metrics applied to early-stage research, and, over the same years, the share of 3M's sales coming from products introduced in the previous five years falling from about a third to about a quarter.[1] One 3M scientist told the magazine plainly: "Invention is by its very nature a disorderly process."[1] The article does not claim Six Sigma alone caused the decline — other explanations were offered too — but it is a documented, widely cited case for a real distinction this module asks you to keep in mind throughout: 5S, standard work, and SMED are built for stabilizing and speeding up a process that is already supposed to produce the same result every time. Pointed at a repeatable production line, they are close to unambiguously good. Pointed at genuinely exploratory work, the same standardizing instinct can suppress the variation that the work actually needs.

The scenario below is a constructed illustration, not a real event.

Almost every Lean tool in this module can also fail for a much more mundane reason: a team runs a three-day kaizen event, moves some racks, times a changeover, feels good about it, and closes the event with a poster showing a bar chart — one bar for "before," one shorter bar for "after." Nobody checked whether the drop was bigger than ordinary day-to-day variation would produce anyway, and nobody checked whether either period was itself stable. Three months later, someone asks whether the change is still in effect, and nobody can say for certain, because the "before" and "after" bars were never anything more than two averages. Everything in this module's tool sections is standard and well documented. The last section, on proving an improvement, is where most real kaizen efforts actually go wrong.

5S and standard work

5S is a workplace organization method in five steps, usually given as Sort (remove what is not needed), Set in order (a place for everything, arranged for the sequence of use), Shine (clean and inspect), Standardize (make the first three a visible, repeatable routine), and Sustain (audit and maintain it).[2] Treated as a one-time cleanup, 5S produces a tidy photo and nothing else; treated as Standardize and Sustain intend, it is a recurring discipline that keeps a workstation in a known state, which is what makes every other tool in this module possible to apply consistently.

Standard work is the companion discipline for the process itself rather than the workstation. The Lean Enterprise Institute defines it as establishing precise procedures for each operator's work in a production process, built from three elements: takt time (the rate at which product must be completed to meet demand), work sequence (the exact order of tasks within that time), and standard in-process stock (the minimum units needed to keep the sequence running smoothly).[3] Standard work is also, quietly, a precondition for everything Modules 16 and 17 teach about control charts: a control chart's "common cause" variation is supposed to be the process's inherent noise, not the noise of three operators each doing the job a different way. Skip standard work, and a chart's control limits partly measure method variation you could have eliminated for free.

Poka-yoke

Poka-yoke (mistake-proofing) is Shigeo Shingo's term for a device or process design that makes a specific human error either physically impossible or immediately self-evident, rather than relying on inspection or attention to catch it afterward.[4] Shingo describes three general mechanisms: a contact method (a physical feature, like an asymmetric locating pin, that only allows correct assembly), a fixed-value method (a counter or a kit that flags a missing or extra step by quantity), and a motion-step method (a sequence sensor that will not let the next step start until the current one is confirmed).[4] This maps directly onto Module 10's PFMEA Detection rating: a poka-yoke that prevents the error is functionally driving Occurrence toward zero, while one that only flags the error after the fact is driving Detection toward 1 (catches it every time) — both lower RPN, but prevention is the stronger fix.

SMED: Single-Minute Exchange of Die

SMED, developed by Shingo largely from his consulting work in Toyota-group plants, is a method for cutting equipment changeover time.[5] Its core move is a single distinction, applied rigorously: separate internal setup (steps that require the machine or furnace to be stopped) from external setup (steps that can be done while it is still running), then (1) convert as much internal work to external as physically possible — staging tooling, pre-heating a fixture, gathering fasteners, all done before the stop — and (2) streamline whatever internal work is left, typically with quick-release fasteners, guided locating features, and standardized settings instead of adjustment. Shingo reports changeovers cut from hours to single-digit minutes at several plants using this method; those specific figures come from his own book rather than an independently audited source, so they are presented here as Shingo's reported results, not independently verified facts.[5] Worked example 1 below applies the same two-step method to a furnace-fixture changeover and checks the result with a hypothesis test rather than taking the improvement on faith.

Kaizen events

Kaizen, from Masaaki Imai's book of the same framing, is continuous, incremental improvement as everyone's ongoing responsibility, not a program run by a separate department.[6] A kaizen event (or kaizen blitz) is a specific, time-boxed application of that idea: a small cross-functional team, a single clearly scoped process, typically three to five days, working from data already collected before the event starts rather than opening with a blank-page brainstorm. That last point is the one teams skip most often. An event that starts with "what do we think is wrong" instead of "here is what the data from the last month shows" tends to converge on whoever argues most persuasively in the room, which is a poor substitute for the root-cause work Module 10 already covers.

Solution selection: the Pugh matrix

Stuart Pugh's concept-selection matrix compares several candidate solutions against one reference concept (the current state, or the most straightforward candidate) across a shared set of criteria.[7] Each candidate is scored, criterion by criterion, relative to the baseline: better (+), the same (S), or worse (−). The scores for a concept sum to a net total, and concepts can be ranked by that total.

net score = (count of +) − (count of −) The baseline concept is defined as scoring S on every criterion against itself, so its own net score is always 0 — every other concept is read as better or worse than that reference point, not against an absolute scale.

Two cautions matter more than the arithmetic. First, the matrix is a facilitation tool for structuring a conversation among people who often disagree, not an algorithm that removes judgment from the decision: a close net score, or a criterion the team privately weights far more heavily than the others, both argue for discussion rather than mechanically picking the top row. Teams that want criteria weighted unequally can multiply each score by a weight before summing (a weighted Pugh matrix); this course uses the simpler unweighted, pure-summation version throughout. Second, a concept can win on a single important criterion and still lose overall once every criterion is counted — Worked example 2 below shows this on real numbers, and it is exactly the situation a matrix is built to surface.

Piloting and proving an improvement with data

A before-and-after comparison built from two averages, or worse, a bar chart with one bar per period, cannot distinguish a real change from ordinary sampling variation. Module 11 covers why in full: even a process with no true change produces two different-looking sample averages essentially every time it is measured twice. A defensible pilot does four things in order. First, pilot the change at a small, contained scale before committing to a full rollout — a single cell, line, or shift, not the whole plant. Second, collect before and after data under comparable conditions, ideally the same measurement system and the same range of normal operating conditions. Third, test the difference with the hypothesis test that matches the data (a two-sample t-test for a measured quantity, a chi-square or two-proportion test for a defect rate) and report an effect size or a practically meaningful number, such as a relative risk or a mean difference, alongside the p-value — a p-value alone says a difference is unlikely to be noise, not that the difference is large enough to matter. Fourth, check each period on its own with a control chart (Modules 16 and 17), because a period average can be pulled by one unusual day without the process's actual level having moved at all; a chart is what tells the two apart. Only after a pilot clears all four does a full-scale rollout, re-confirmed the same way, belong on the table — which is exactly what the capstone (Module 19) will do at full scale, building on the smaller pilot this module's Worked example 3 pilots first.

Worked examples

Worked example 1: furnace-fixture changeover time, before and after SMED

The data below is a constructed example, not a real production run. Setting: time in minutes to change the braze-furnace fixture between part numbers, timed to the nearest tenth of a minute, 16 changeovers before and 16 after separating internal from external setup steps (staging the next fixture and pre-heating it externally, and replacing a bolted mount with a quick-release clamp for the one step that has to happen with the furnace stopped).

Table 1. Furnace-fixture changeover time (min), 16 changeovers before and 16 after (constructed data).
Group12345678910111213141516
Before73.670.380.373.465.076.789.084.362.955.563.972.541.869.255.862.5
After23.725.129.533.326.235.223.029.132.427.622.521.524.328.320.925.7

Mean before 68.54 min (s = 11.80), mean after 26.77 min (s = 4.29), a difference of 41.78 min. Before testing, check the equal-variance assumption behind the standard pooled t-test: Levene's test gives p = 0.0084, well under 0.05, so the two periods' variances are not equal (the after period is both faster and more consistent, which makes sense — a quick-release clamp removes a source of adjustment-to-adjustment variation along with time). The textbook response is to prefer Welch's t-test, which does not assume equal variances: t = 13.31 on 18.9 degrees of freedom (Welch's own adjusted df, versus 30 for the pooled test), 95 % CI on the difference 35.20 to 48.35 min, p < 0.0001. Cohen's d = 4.70, an enormous effect size by any convention. Here the equal-variance violation changes which test is technically correct but not the practical conclusion, because the effect is so large that both the pooled test (p < 0.0001) and Welch's test agree; for a smaller, more marginal effect, checking this assumption before trusting the p-value can matter a great deal more than it does here.

Hypothesis tests calculator

Worked example 2: Pugh matrix, a furnace-fixture redesign

Constructed example, not a real selection workshop.

Module 10's root-cause work traced leak-test rejects partly to furnace-braze joint gap. A solution-selection workshop scores four fixture-redesign concepts against the current pin fixture on five criteria: how much the concept improves joint-gap repeatability, its effect on changeover time, its unit cost, operator ergonomics, and maintainability.

Table 2. Pugh matrix, furnace-fixture concepts vs. the baseline pin fixture (constructed data). + better than baseline, S same, − worse than baseline.
ConceptGap repeatabilityChangeover timeUnit costErgonomicsMaintainability+SNet
Baseline (current pin fixture)SSSSS0050
Machined locating block++S+311+2
Spring-loaded clamp+++320+1
Modular quick-change++++410+3
Ceramic insert+S131-2

The modular quick-change concept wins on net score (+3), and notice what that number is doing: it is not the cheapest option (both it and the spring-loaded clamp cost more than baseline; the machined locating block is actually the only one that is cheaper), and it is not the only concept that improves gap repeatability (all four do — that was the point of the workshop). It wins because it is the only concept that improves on four of the five criteria at once, trading a higher unit cost for gains everywhere else. The ceramic insert is the instructive failure case: it does improve the one criterion the team most cared about (gap repeatability, since ceramic does not thermally distort at furnace temperature), but it nets to −2 once cost, ergonomics, and maintainability (brittle, harder to inspect, more care needed in handling) are all counted. A team that stopped at "does it fix the gap problem?" would have picked a net loser. This is not a reason to stop asking that question — it is a reason to also ask the other four.

Worked example 3: piloting the modular quick-change fixture

The data below is a constructed example, not a real production run. Setting: brazed aluminium heat-exchanger cores, 100 % helium leak test at final inspection, one lot per shift. The modular quick-change fixture from Worked example 2 was piloted on one cell for 14 shifts, compared against 15 shifts on that same cell immediately before the change. (The full-scale, 30-shift-per-period confirmation after a plant-wide rollout is the capstone, Module 19.)

Fifteen lots before the change ran 0.0441 (4.41 %) leakers on average, close to the Module 3 charter's 4.2 % baseline, over an average lot size of 105.7 cores. Fourteen lots after ran 0.0201 (2.01 %), over a smaller average lot size of 71.1 cores — smaller because a pilot cell runs less volume than the full line. Both periods, charted individually as p charts, show no rule violation: neither is drifting or unstable on its own. Notice something else about both charts, though: at these lot sizes and this rate, the 3-sigma lower control limit works out to 0.0 on every single point in both periods, simply because the rate is low enough that "mean minus three sigma" goes negative and gets floored at zero. That means the p chart's own 3-sigma rule cannot, by itself, flag the after period as a special-cause improvement over the before period — a chart of the after period alone would look identical whether the fixture helped or not. This is exactly why step three of piloting (above) calls for a hypothesis test in addition to the charts, not instead of them: the charts confirm each period is internally stable; the test is what actually compares the two periods to each other.

Table 3. Leak-test rejects by lot, pilot cell, before and after the fixture change (constructed data).
PeriodLotsTotal coresTotal leakersRate
Before151586704.41 %
After14996202.01 %

Treating this as a 2×2 table (leak or pass, before or after) and running a chi-square test of independence — equivalent to a two-proportion test — gives χ² = 10.524 on 1 degree of freedom, p = 0.0012: comfortably below 0.05, so the drop is not plausibly just sampling noise. Alongside the p-value, the practical-significance numbers: relative risk (after ÷ before) = 0.455, meaning the pilot cell's leak rate is a little under half what it was, and a risk difference of -2.41 percentage points. Cramér's V, a chi-square effect-size measure, comes out small (0.064) — worth flagging as a common trap on its own: V is small here mainly because leaks are rare in both periods (most of both tables is "Pass"), not because the change was unimportant. For a rare-event table like this one, relative risk and risk difference are the practically meaningful numbers; a small Cramér's V should not be read as "the effect did not matter." The pilot supports rolling the fixture out further, with the expectation — not yet the proof — that a full-scale, longer-running confirmation will show at least as large an improvement; Module 19 shows what that confirmation actually found.

Pilot before, p chart (control chart builder, inputs collapsed)

Pilot after, p chart (control chart builder, inputs collapsed)

Common mistakes

Exercises

Exercise 1: pallet-changeover time on a machining cell

Constructed example, not a real production run. Setting: time (min) to change the fixture pallet on a CNC machining cell, 10 changeovers before and 10 after fitting a locating pin that only accepts the pallet one way round (a contact-method poka-yoke that also happens to speed up the changeover, since the operator no longer has to check alignment by eye).

Table 4. Pallet-changeover time (min), 10 before and 10 after (constructed data).
Group12345678910
Before11.012.712.616.68.717.214.221.621.823.6
After10.98.911.112.87.410.58.912.66.98.2

Tasks. (a) Compute both group means and standard deviations. (b) Run a two-sample t-test (pooled, equal-variance assumption) and report t, degrees of freedom, and the p-value. (c) Compute Cohen's d and compare its size to Worked example 1's. Why would you expect this pilot's effect size to be smaller than a full SMED redesign's?

Show the worked solution

(a) Mean before 16.00 min (s = 5.04), mean after 9.82 min (s = 2.07), difference 6.18 min.

(b) t = 3.59 on 18 degrees of freedom, p = 0.0021, 95 % CI on the difference 2.56 to 9.80 min. Still comfortably significant, but nowhere near Worked example 1's p < 0.0001 — this one genuinely needed the t-table, not just a glance at the numbers.

(c) Cohen's d = 1.60, still a large effect by conventional benchmarks, but well under Worked example 1's 4.70. A single locating pin removes one specific source of delay (checking and correcting alignment); a full SMED redesign that separates internal from external setup and redesigns the clamping attacks every source of delay in the changeover at once, so a bigger effect size from the fuller redesign is the expected pattern, not a coincidence.

Exercise 2: Pugh matrix for a connector poka-yoke

Constructed example, not a real selection workshop.

A team is choosing a device to stop an operator inserting a wiring connector backwards, scored against the current baseline (visual inspection only) on four criteria: how well it prevents the error, unit cost, its impact on cycle time, and the training burden it adds.

Table 5. Pugh matrix, connector poka-yoke concepts vs. visual inspection only (constructed data).
ConceptError preventionCostCycle-time impactTraining burden
Baseline (visual inspection only)SSSS
Mechanical stop pin++S+
Optical sensor interlock+S
Asymmetric locating feature+S+S

Tasks. (a) Compute the net score for each concept. (b) Which concept wins? (c) The optical sensor interlock has the same error-prevention score as the two winning concepts. Why does it still net out worst of the three real alternatives?

Show the worked solution

(a) Mechanical stop pin: net = 3. Optical sensor interlock: net = -1. Asymmetric locating feature: net = 2.

(b) The mechanical stop pin wins, with a net score of +3, beating every other concept including the baseline.

(c) All three alternatives improve error prevention equally (each scores + there), so that criterion does not distinguish them at all — the decision is entirely made by the other three criteria. The optical sensor costs more and slows the cycle down (two minuses), while the stop pin and the locating feature cost the same or less than baseline and do not hurt cycle time. Looking only at error prevention, all three would look identical; the matrix's value here is entirely in forcing the other criteria onto the same page.

Quiz

Ten questions. Score 70 % or more to mark the module complete on this device.

1. In 5S, "Standardize" and "Sustain" matter most because
2. Standard work's three elements are
3. A poka-yoke that makes an assembly error physically impossible, versus one that only signals the error after it happens, mainly differs in that the first
4. SMED's core method is to
5. In Worked example 1, Levene's test showing p = 0.0084 led to preferring Welch's t-test over the pooled t-test because
6. A kaizen event that opens with a blank-page brainstorm instead of data already collected beforehand mainly risks
7. In a Pugh matrix, a concept's net score is computed as
8. In Worked example 2, the ceramic insert concept improved gap repeatability but still had the worst net score because
9. In Worked example 3, the p chart's lower control limit was 0.0 for every point in both the before and after periods because
10. In Worked example 3, Cramér's V came out small (0.064) even though the leak rate roughly halved, mainly because
Answer key
  1. b. Without them, 5S is a one-time cleanup, not a recurring discipline.
  2. c. Takt time, work sequence, standard in-process stock.
  3. a. Prevention drives Occurrence toward zero, not just Detection.
  4. d. Separate internal/external, convert internal to external, streamline what remains.
  5. b. The variances were shown unequal, which the pooled test assumes they are not.
  6. c. Converges on the most persuasive argument, not the actual cause.
  7. d. Count of pluses minus count of minuses.
  8. a. It scored worse on three other criteria, outweighing the one strong criterion.
  9. c. Mean minus three sigma is negative at this rate and lot size, floored at zero.
  10. b. Leaks were rare in both periods, which keeps this measure small regardless of practical importance.

Key takeaways

References

All web sources accessed 2026-09-10 unless noted.

  1. Hindo, B. "At 3M, a struggle between efficiency and creativity." BusinessWeek, 11 June 2007. https://effectuation.org/hubfs/Journal%20Articles/2016/06/3m-struggle-between-efficiency-and-creativity.pdf
  2. Hirano, H. (1995). 5 Pillars of the Visual Workplace. Productivity Press. https://www.routledge.com/5-Pillars-of-the-Visual-Workplace/Hirano/p/book/9781563270475 (catalogue-level; the five steps)
  3. Lean Enterprise Institute. Lexicon: "Standardized Work." https://www.lean.org/lexicon-terms/standardized-work/
  4. Shingo, S. (1986). Zero Quality Control: Source Inspection and the Poka-yoke System. Productivity Press. https://www.routledge.com/Zero-Quality-Control-Source-Inspection-and-the-Poka-Yoke-System/Shingo/p/book/9780915299072 (catalogue-level; the three mechanisms)
  5. Shingo, S. (1985). A Revolution in Manufacturing: The SMED System. Productivity Press. https://www.routledge.com/A-Revolution-in-Manufacturing-The-SMED-System/Dillon-Shingo/p/book/9780915299034 (catalogue-level; the internal/external distinction is confirmed at this depth, specific numeric changeover-time claims are Shingo's own reported figures, not independently audited, and are presented as such)
  6. Imai, M. (1986). Kaizen: The Key to Japan's Competitive Success. McGraw-Hill. https://archive.org/details/kaizen00masa (catalogue-level)
  7. Pugh, S. (1991). Total Design: Integrated Methods for Successful Product Engineering. Addison-Wesley. https://en.wikipedia.org/wiki/Stuart_Pugh (secondary: the concept-selection matrix method, via a summary source)