Four complete A3 problem-solving case studies across four industries — manufacturing, software, healthcare, office — with every section filled in using realistic numbers. Each follows the canonical Toyota 7-section layout: Background · Current Condition · Goal · Root Cause · Countermeasures · Implementation Plan · Follow-up. Read these as templates for your own A3, not as theory. The A3 discipline is the one-page constraint: if your thinking does not fit, your thinking is not done.
The four case studies at a glance
| Industry | Problem | Root cause | Result after follow-up |
|---|---|---|---|
| Manufacturing | Fastener cell cycle-time 18% over takt | Tool-changeover variance + non-standard work | +14% throughput, on-takt |
| Software / SRE | On-call paging volume causing engineer burnout | Alert thresholds set on raw metrics, not user impact | Pages down 60%, sleep-disrupting pages near zero |
| Healthcare | Emergency Dept door-to-doc time 47 min | Triage bottleneck at single nurse station | Door-to-doc 25 min (-22 min, -47%) |
| Office / Finance | Re-bill rate at 8% (~$340k/quarter rework cost) | No structured handoff between sales and billing | Re-bill rate 2%, ~$255k/quarter saved |
Each case study below is written as if you were reading the actual A3 sheet, section by section. Numbers and details are realistic but composite — the patterns reflect dozens of real Lean engagements, not any single client.
Case study 1 — Manufacturing: Fastener cell cycle-time reduction
1. Background
Cell C-3 produces 14 SKUs of automotive fasteners (M6–M12 hex bolts and flanged screws) for a Tier-2 customer with weekly nominated quantities. The cell runs two shifts, five days, with a target takt of 38 seconds per part. Over the last quarter, cell C-3 has consistently run 18% over takt — observed actual 45 seconds — forcing weekend overtime to meet weekly commitment. Customer scorecard has flagged on-time-in-full at 91% versus the 98% target. Continued overrun risks losing 30% of the next-year contract.
2. Current condition
- Actual cycle time: 45 seconds (mean across 4 weeks, n = 8,420 parts logged via MES)
- Takt: 38 seconds (customer quantity / available run time)
- Overrun: +7 seconds = +18%
- Tool changeovers per shift: 4 average; 18 minutes mean changeover, σ = 7 minutes (high variance)
- Operator-to-operator variance: 39 to 52 seconds depending on operator (n = 6 operators tracked)
- Standard work compliance audit: 60% (most operators using personalised technique on 3 of 7 process steps)
3. Goal / Target
Cycle time at or below 38 seconds (on takt) within 6 weeks, sustained for 4 consecutive weeks, with operator-to-operator variance under ±5%. On-time-in-full back to 98%.
4. Root cause analysis
Fishbone first to map the landscape (Methods, Machines, Materials, Manpower, Measurement, Environment), then 5 Whys on the dominant Methods branch.
- Why is cycle time over takt? Operators take 6–13 extra seconds on the inspection-and-load step.
- Why does inspection take that long? Each operator is checking different features in different orders.
- Why is the inspection step not standardised? The original standard work sheet was written in 2019 for a previous SKU mix and was never updated when the M10 flanged variants were added.
- Why was it not updated? No process owner assigned for standard work maintenance after the 2022 reorganisation.
- Why is there no process owner? Standard work was treated as a one-time deliverable rather than a living document — ownership was assumed but never written down.
Verified root cause: No process owner for standard work maintenance, leading to drift and operator-specific technique.
Secondary root cause: Tool-changeover variance (σ = 7 min) reflects no standard quick-changeover sequence — SMED principles never applied to this cell.
5. Countermeasures
- Reissue standard work sheets for all 14 SKUs with current process, validated by 2 senior operators and the manufacturing engineer.
- Assign cell production engineer as standing process owner with quarterly review cadence.
- SMED workshop on tool changeover — target reduce mean from 18 to 9 minutes, σ from 7 to 2.
- Daily start-of-shift standard-work confirmation by team leader (5 minutes, observation-based).
- Weekly cycle-time review at cell huddle with control chart on the cell-side board.
6. Implementation plan
| Action | Owner | By | Verify |
|---|---|---|---|
| Reissue standard work sheets | Cell engineer | Week 1 | 2-operator + ME sign-off |
| Process-owner assignment posted | Plant manager | Week 1 | Email + cell board |
| SMED workshop run | Lean lead | Week 2 | Before/after stopwatch study |
| Operator retraining (all 6) | Cell engineer | Week 3 | Skills matrix updated |
| Daily standard-work confirmation | Team leader | Week 3 onward | Daily checklist |
| Cycle-time control chart live | Cell engineer | Week 4 | Board photo weekly |
7. Follow-up
Week 6 measurement (n = 7,860 parts):
- Mean cycle time: 37 seconds (target ≤38) ✅
- Operator variance: ±4% (target ±5%) ✅
- Tool changeover mean: 10 minutes, σ = 2.5 (target 9 min, σ 2) — slightly above goal but acceptable trend
- On-time-in-full last 2 weeks: 97% (target 98%, trending upward)
- Throughput: +14% versus baseline
Sustaining actions: quarterly standard-work review by process owner; cycle-time control chart added to plant tier-2 daily metrics. Re-audit at month 3.
Case study 2 — Software: On-call paging burnout
1. Background
The platform SRE team (8 engineers) rotates on-call for 12 production services (~140 alerts configured). In the last quarter, two engineers left the company citing on-call burden, exit interviews specifically called out the volume of after-hours pages. Engineering survey scores on "on-call is sustainable" dropped from 6.2 to 3.4 (out of 10). Hiring backfill is 4–5 months and the remaining 6 engineers are at risk of cascading burnout.
2. Current condition
- Pages per week, team total: 142 mean (last 8 weeks, PagerDuty data)
- Pages per engineer per week: ~17.7
- Sleep-disrupting pages (00:00–06:00 local): 38 per week, ~4.7 per engineer
- Pages requiring engineer action (vs. self-resolving / informational): 31% — meaning 69% of pages are noise
- Alert-to-incident conversion: 12% (88% of pages do not become real incidents)
- Engineer survey “on-call is sustainable”: 3.4 / 10 (down from 6.2)
3. Goal / Target
Reduce total pages per week to under 60 (a ~60% reduction), reduce sleep-disrupting pages to under 1 per engineer per week, and raise the "on-call is sustainable" survey score to ≥6 by week 8. Stop attrition driven by on-call burden.
4. Root cause analysis
Fishbone across Alerts, Services, Tooling, Process, then 5 Whys on the dominant Alerts branch.
- Why are pages so high-volume? Most alerts trigger on raw infrastructure metrics (CPU, memory, queue depth) regardless of user impact.
- Why use raw metrics? Alerts were authored 4 years ago when each service had its own owning engineer who knew which thresholds mattered.
- Why are those alerts still in use? No one reviews them; alerts are treated as "set once and forget" configuration.
- Why is there no review process? No team-level ownership of the alerting policy — each service is whoever-last-touched-it.
- Why is there no policy? The team grew from 3 to 8 engineers without ever formalising on-call standards; original tribal knowledge faded as people left.
Verified root cause: No team-level alerting policy or review cadence. Alerts proliferate based on infrastructure metrics rather than user-visible symptoms.
5. Countermeasures
- Adopt SLO-based alerting — page only when a user-visible SLO (latency, error rate, availability) is at risk of burn-rate violation. Move infrastructure-metric alerts to ticket queue, not pager.
- Quarterly alert review — team reviews every paging alert, kills the ones that did not produce a real incident response in the last 90 days.
- On-call retrospective at end of every rotation — departing on-call engineer reviews every page with the team, classifies as actionable / noise / informational, kills or downgrades the noise ones immediately.
- Auto-resolve for self-healing alerts — if the underlying signal returns to normal within 5 minutes, suppress the page entirely.
- Document on-call standards — one-page team policy: severity definitions, escalation rules, what should and should not page.
6. Implementation plan
| Action | Owner | By | Verify |
|---|---|---|---|
| SLO catalogue for top 5 services | Senior SRE | Week 2 | Approved by service owners |
| Migrate top 5 services to SLO alerting | SRE rota | Week 4 | Old alerts disabled, new alerts firing tested |
| End-of-rotation retrospective process | Team lead | Week 1 (every rotation) | Doc + alert kill log |
| Auto-resolve config for noisy alerts | Tooling SRE | Week 3 | PagerDuty rule deployed |
| On-call standards doc published | Team lead | Week 2 | Linked from runbook wiki |
| Repeat for remaining 7 services | SRE rota | Week 8 | All 12 services SLO-based |
7. Follow-up
Week 8 measurement:
- Pages per week, team total: 52 (target <60) ✅ — down from 142, a 63% reduction
- Sleep-disrupting pages per engineer per week: 0.6 (target <1) ✅
- Alert-to-incident conversion: 54% (up from 12% — remaining alerts are mostly real)
- Survey "on-call is sustainable": 6.8 / 10 (target ≥6) ✅
- Attrition since start: 0 (vs. 2 in prior quarter)
Sustaining actions: end-of-rotation retro is now standard process; quarterly alert review on the team calendar; on-call survey re-run quarterly.
Case study 3 — Healthcare: ED door-to-doc time
1. Background
The Emergency Department of a 350-bed urban hospital sees 220 patients/day on average. Door-to-doc time (the interval from patient registration to first physician contact) currently averages 47 minutes — well above the 30-minute internal target and the published clinical-quality benchmark for an urban ED of this acuity mix. Patient satisfaction has dropped, left-without-being-seen rate climbed to 4.5%, and three serious-incident reviews in Q1 cited triage-stage delays as a contributing factor.
2. Current condition
- Mean door-to-doc: 47 minutes (n = 6,580 patients, last 4 weeks, EHR triage timestamps)
- 90th percentile: 92 minutes
- Triage station throughput: 1 nurse station, mean 8 minutes per patient (full vitals + ESI assessment)
- Wait between registration and triage: mean 18 minutes (single longest queue point)
- Wait between triage and physician: mean 21 minutes
- Left-without-being-seen rate: 4.5% (clinical target ≤2%)
- Patient satisfaction (HCAHPS-equivalent ED survey): 64% (down from 78% one year ago)
3. Goal / Target
Mean door-to-doc time at or below 30 minutes within 10 weeks, sustained for 4 consecutive weeks. Left-without-being-seen rate at or below 2%. No triage-stage delays cited in incident reviews.
4. Root cause analysis
5 Whys on the triage queue (the dominant constraint visible in the value-stream walk).
- Why is door-to-doc 47 minutes? 18 min queueing for triage + 8 min triage + 21 min waiting for doc = 47.
- Why is the triage queue 18 minutes? Single triage nurse station serves all arrivals; arrival rate exceeds throughput during peak hours (10:00–22:00).
- Why only one station? Original ED layout from 2008 had one triage room; renovation in 2019 added a second consultation room but it is currently used as overflow exam space.
- Why is the second room not used for triage? No standard process for opening a second triage station; nurse staffing pattern assumes single-station throughput.
- Why is staffing built around single-station throughput? Staffing model dates to 2018 patient-volume baseline; daily census has grown 30% since then but the staffing pattern was never re-baselined.
Verified root cause: Triage capacity has not scaled with arrival rate; staffing and physical-space allocation reflect a 30%-lower baseline.
5. Countermeasures
- Open the second triage station during peak hours (10:00–22:00) using the renovated consultation room.
- Re-baseline nurse staffing against current census; add one nurse to the peak shift to staff the second station.
- Standard triage protocol — agreed ESI assessment sequence, target 6 minutes per patient (down from 8) without losing clinical rigor.
- Pull-trigger for opening second station — if waiting room exceeds 5 patients waiting for triage, charge nurse opens the second station immediately.
- Daily metrics board at nurse station with door-to-doc, LWBS, and patients-per-shift, updated each shift change.
6. Implementation plan
| Action | Owner | By | Verify |
|---|---|---|---|
| Re-baseline staffing model | ED nurse manager + HR | Week 2 | Updated FTE plan signed |
| Renovate consult room → triage 2 | Facilities | Week 3 | Walkthrough sign-off |
| Standard triage protocol drafted + trained | Lead RN | Week 4 | All triage RNs signed off |
| Pull-trigger rule live | Charge nurses | Week 5 | Posted at nurse station |
| Second station staffed peak hours | Nurse scheduler | Week 6 | Schedule live, 4 weeks consistent |
| Daily metrics board | Charge nurse rota | Week 6 onward | Photo audit weekly |
7. Follow-up
Week 10 measurement (n = 7,140 patients, weeks 7–10):
- Mean door-to-doc: 25 minutes (target ≤30) ✅ — down 22 minutes / 47%
- 90th percentile: 48 minutes (down from 92)
- Triage queue mean wait: 6 minutes (down from 18)
- Left-without-being-seen rate: 1.8% (target ≤2%) ✅
- Patient satisfaction: 74% (up from 64%, target 78% over next 2 quarters)
- Triage-stage delays in incident reviews this period: 0
Sustaining actions: staffing model re-baseline annually against rolling 12-month census; daily metrics board permanent; LWBS reviewed weekly at ED leadership huddle.
Case study 4 — Office / Finance: Billing accuracy and re-bill rate
1. Background
The B2B services division of a mid-market firm bills 1,800 invoices per quarter (~$28M revenue). Currently 8% of invoices require a re-bill (correction and resend) due to errors in customer details, line items, contract reference, or pricing. Each re-bill costs an estimated $235 in finance team time and delays cash collection by an average of 11 days. CFO estimates the annualised cost of re-billing at ~$1.4M, plus 14 days of working-capital drag.
2. Current condition
- Re-bill rate: 8.0% (last 4 quarters, n = 7,210 invoices, AR ledger)
- Top 3 error categories: wrong contract reference (32%), wrong line-item pricing (28%), wrong customer entity (21%); other 19%
- Mean cost per re-bill: $235 (1.5 hours of senior billing analyst + 0.5 hour controller review)
- DSO impact: +11 days mean delay on re-billed invoices vs. clean ones
- Quarterly cost of re-billing: ~$340k
- Number of handoffs from sale closure to invoice issued: 4 (Sales rep → Sales ops → Billing → Controller)
- Errors introduced at each handoff (estimated): ~30% at the sales-rep-to-sales-ops boundary
3. Goal / Target
Re-bill rate at or below 2% within 12 weeks, sustained over 2 consecutive quarters. Estimated annualised cost reduction: ~$1.0M. Reduce DSO impact by 8 days on average.
4. Root cause analysis
Fishbone first across People, Process, Information, Tooling, then 5 Whys on the Information branch (the dominant cluster).
- Why is the re-bill rate 8%? Most errors enter at the sales-rep-to-sales-ops handoff — mismatched contract references and pricing.
- Why do errors enter there? Sales reps email a free-form deal summary to sales ops, who then transcribe into the billing system.
- Why is it free-form? No standard handoff template exists; each rep formats the summary differently.
- Why no template? Billing was never involved in the deal-handoff design; the sales process treats "billing" as downstream and not part of the closing workflow.
- Why is billing not part of the closing workflow? Organisational silo between Sales and Finance — no joint process owner for the order-to-cash handoff.
Verified root cause: No standard handoff between Sales and Billing, with no joint process owner. Errors are introduced at a free-form text boundary.
5. Countermeasures
- Standard order-to-cash handoff form in CRM — structured fields for contract reference, customer entity, line-item pricing, billing schedule. Sales rep cannot mark deal "closed-won" without completing it.
- Validation rules in CRM — contract reference must match an existing master-contract record; pricing must match contract pricing tables.
- Joint process owner — named Sales Ops + Billing co-leads for the order-to-cash handoff, monthly review.
- Daily exception report — any handoff form filed with missing/invalid fields routed to Sales Ops for fix before invoice issuance.
- Quarterly retrospective on top error categories with both Sales and Finance present.
6. Implementation plan
| Action | Owner | By | Verify |
|---|---|---|---|
| Handoff form designed (Sales + Billing) | Joint process owners | Week 3 | Approved by VP Sales + CFO |
| CRM custom object built | RevOps engineer | Week 5 | UAT pass |
| Validation rules deployed | RevOps engineer | Week 6 | Test with 20 sample deals |
| Sales rep training (cohort of 24) | Sales enablement | Week 7 | 100% completed + signed off |
| Daily exception report live | Billing analyst | Week 8 | Daily distribution + tracking log |
| Process owner monthly review | Joint owners | Week 8 onward | Meeting notes + action log |
7. Follow-up
Week 12 measurement (Q-end, n = 1,810 invoices):
- Re-bill rate: 2.1% (target ≤2%) — narrowly missed target but at directional goal ✅
- Top error category: wrong contract reference dropped from 32% of errors to 9%
- Mean cost of re-billing: $245/invoice (similar — but ~75% fewer re-bills)
- Quarterly cost of re-billing: ~$85k (down from $340k, saving ~$255k)
- DSO impact reduced: -7 days mean on what would have been re-billed (target -8)
- Annualised cost saving: ~$1.0M in finance team time + working-capital release
Sustaining actions: validation rules in CRM are now part of the deal-closure workflow; quarterly process-owner review continues. Next A3 cycle targeting the remaining 19% of error categories that the structured handoff did not address.
Common patterns across the four cases
Reading the four side-by-side reveals patterns that show up in nearly every A3, regardless of industry.
1. The current condition section is where most A3s fail
In all four examples, the current condition has at least 5 specific numbers with citation context (sample size, period, source). When a current condition reads "cycle time is too long" without a number, the rest of the A3 is built on sand. The discipline is to count, time, and measure before drafting the goal box.
2. Root cause hides one level deeper than you expect
The first three Whys are usually obvious to the team. The 4th and 5th Why — where ownership, organisational structure, or process design is the real cause — is where the actionable countermeasure lives. In the manufacturing case, "tool changeover takes too long" is the symptom; "no process owner for standard work" is the root cause. In the office case, "invoices have errors" is the symptom; "no joint process owner for order-to-cash" is the root cause.
3. Countermeasures should map one-to-one to root cause lines
Look at any of the four countermeasure lists — each numbered item directly addresses a verified root-cause finding. None of them are "train the team better" or "communicate more". Vague countermeasures are a sign that the root-cause work was not finished.
4. Implementation plans need owner, deadline, and verification — not just a deadline
The 4-column implementation tables (action / owner / by / verify) appear identical across cases. The verify column is what closes the loop — without it, "done" means "an email was sent" rather than "the change is in production and measured".
5. Follow-up uses the same units as current condition
Notice how each follow-up section measures the same metric, sample, and source as the current condition. This is non-negotiable: if your current condition is in minutes and your follow-up is in "significantly faster," you have not measured anything. The A3 closes when the metric you started with reaches the target you set.
Run the cause analysis live
Three of these four cases used 5 Whys or Fishbone (or both) in the root-cause box. Use the free tools to do the analysis interactively, then paste the result onto your A3 template.
Open the 5 Whys tool →Common pitfalls these case studies avoid
- Solution-jumping. None of the four jumped to a countermeasure before completing root-cause work. The temptation to write "add staff" in the manufacturing case before discovering the standard-work ownership gap is exactly the trap A3 is designed to prevent.
- Blaming people. All four root-causes are process root causes, not person root causes. "The operator was slow" or "the sales rep made errors" are dead ends — the actionable fix is the system around the person.
- Missing baseline. Every case has a numbered current condition before any goal is set. The gap between current and goal is what determines countermeasure ambition.
- Goals without timelines. Every goal box has a date. "Reduce cycle time" without "by week 6" is a wish, not a target.
- Follow-up that does not measure. Every follow-up section returns to the same units as current condition. Anecdotes do not close A3s.
Common questions
How do I scale this to my own A3?
Start by downloading the free A3 template. Then walk the 7 sections in order: do not write the countermeasure section until current condition and root cause are filled in with real data. The discipline is sequential — each section depends on the one before it. Treat the template as a forcing function, not a fill-in-the-blanks form.
What if my problem touches more than one team?
Cross-functional A3s are common — the office/billing case here is one. The convention is that the A3 has a single named author (the person closest to the problem) plus a coach (often a senior leader from a different function). Other functions appear in the implementation plan as named owners. The author owns the page; ownership of countermeasures distributes to the right teams.
Can I use A3 for a problem I already know the answer to?
You can, but the value drops considerably. A3 is most useful when the team thinks they know the answer but has not verified the root cause — the discipline forces them to slow down. If the answer is genuinely obvious and the data supports it, an A3 may be overkill; a one-page change request or kanban card is enough. Reserve A3 for problems where the cost of a wrong answer (or a too-shallow answer) is high.
How do these case studies relate to 8D?
All four are internal A3s — no customer involved, no external corrective-action obligation. If any of these problems hit a Tier-1 customer (say the manufacturing fastener-cell defects went out the door), the team would do the thinking on an A3 and then translate the finished thinking into the customer's 8D template for submission. See A3 vs 8D for the hybrid pattern.
What software should I use for digital A3s?
The simplest option is the downloadable A3 template in Excel, Word, or PDF. For collaborative real-time work, Miro, FigJam, and Mural all have A3-style templates. PowerPoint and Google Slides also work fine — the constraint is the one-page format, not the tool. Avoid tools that encourage scrolling; the visual gestalt of seeing the whole story at once is the point.
Do A3 case studies need to be this long?
The actual A3 sheet is one page — what you see here is the "A3 read aloud" with each section expanded for explanatory purposes. A real A3 in production fits all 7 sections onto a single landscape A3 sheet (297 × 420 mm). The verbosity here is for teaching; the discipline in practice is the opposite — severe compression onto one page.
What to read next
- A3 problem solving — the complete Toyota guide — full method, history, coaching loop, and worked Lean example.
- Free A3 template (Excel, Word, PDF) — the canonical Toyota 7-section layout to start your own.
- A3 vs 8D — when to use the internal A3 vs. the customer-facing 8D, and the hybrid pattern.
- 8D report examples — five customer-facing 8Ds across automotive, aerospace, medical-device, food, and electronics.
- PDCA cycle — the Plan-Do-Check-Act backbone that every A3 traces.
- 5 Whys library — the iterative questioning technique used in the root-cause box of three of these four cases.
- Fishbone diagram — the cause-mapping technique used to identify the dominant branch before 5 Whys.